Visual Credit Audit Reveals Overestimated Spatial Reasoning in AI
A new method called Visual Credit Audit (VCA) exposes that many correct answers on spatial reasoning benchmarks for multimodal large language models (MLLMs) do not actually rely on visual information. The study, published on arXiv (2607.27069), shows that 12.73-26.25% of correct decisions are uncredited to the image, meaning the model could answer correctly without seeing it. VCA compares model performance with text-only and blank controls, and uses image permutation to measure dependence on visual evidence. Across four open MLLMs and two spatial benchmarks, matched image permutation reduced dependence-credited correctness (D-CC) by 21.25-47.80 points, with all 95% confidence intervals above zero. The method is training- and label-free for the first audit, and applying labels yields D-CC. The findings suggest that current benchmarks may overestimate visual reasoning capabilities, as models often rely on language priors or spurious correlations.
Key facts
- Visual Credit Audit (VCA) separates whether the image supports the model's decision more than text-only and blank controls.
- 12.73-26.25% of correct decisions are uncredited to the image across four open MLLMs and two spatial benchmarks.
- Matched same-split image permutation reduces D-CC by 21.25-47.80 points.
- All paired 95% intervals above zero indicate significant visual dependence.
- The first audit is training- and label-free and does not require an answer flip.
- Applying labels yields dependence-credited correctness (D-CC).
- Prediction alignment extends the audit to errors.
- The study uses fixed-pixel relation contrasts and a 3x3 evidence-source factor.
Entities
Institutions
- arXiv