TRAPSBench: VLMs Perceive Answerability but Fail to Abstain
A recent investigation published on arXiv (paper 2608.13167) indicates that Vision-Language Models (VLMs) can recognize when they do not have enough visual evidence to respond to a query, yet they do not communicate this uncertainty by choosing to abstain. The study introduces TRAPSBench, a video benchmark with 1,404 physics pairs, where a specific alteration renders the outcome impossible to determine visually. Additionally, the researchers present a novel metric called Penalized Epistemic Calibration Score (PECS), which incentivizes models for correct answers when outcomes are clear and for abstaining when they are not. Among 16 VLMs from five families, the ability to self-restrain is inadequate, with the highest PECS score being 0.292. The challenge seems to lie in expression rather than perception, as linear probes can effectively decode answerability from hidden states with an AUROC of up to 0.91 across physics domains. The results are consistent across various fields, indicating that while VLMs capture epistemic states, they lack the ability to act upon them. This research is significant for the safe application of AI in critical scenarios where erroneous responses could lead to severe repercussions.
Key facts
- TRAPSBench is a procedurally generated video benchmark with 1,404 matched physics pairs.
- Penalized Epistemic Calibration Score (PECS) is introduced as a new metric.
- 16 VLMs across five families were evaluated.
- Best PECS score achieved was 0.292.
- Linear probes decode answerability from hidden states with up to 0.91 AUROC.
- Steering a single-layer void direction causally induces or suppresses abstention.
- The paper is available on arXiv with ID 2608.13167.
- The study shows VLMs can internally distinguish when abstention is required but fail to express it.
Entities
Institutions
- arXiv