Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
A recent paper published on arXiv (2608.06938) disputes the common belief that the shortcomings in visual counter-commonsense reasoning of multimodal large language models (MLLMs) are due to a lack of visual grounding. The authors' empirical findings indicate that MLLMs already possess the necessary visual information and that the correct answers are present within their decoding framework. The actual limitation lies in the shared language decoder, which tends to prioritize dominant language priors, particularly in low-frequency factual contexts. To tackle this issue, they introduce a text-anchored data construction method, featuring Fact-Frequency Distillation (FFD) as a key element, which assesses the strength of commonsense facts and refines verifiable counter-commonsense instances. The full paper can be accessed at https://arxiv.org/abs/2608.06938.
Key facts
- The paper is titled 'Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning'.
- It is published on arXiv with identifier 2608.06938.
- The research focuses on visual counter-commonsense reasoning in multimodal large language models (MLLMs).
- The authors' empirical analysis shows that MLLMs already capture relevant visual evidence and the correct answer exists in their decoding space.
- The bottleneck is identified as the shared language decoder, which favors dominant language priors over visual evidence.
- The problem is more pronounced for low-frequency factual scenarios.
- The paper proposes a text-anchored data construction pipeline.
- The core component of the pipeline is Fact-Frequency Distillation (FFD), which estimates the prior strength of commonsense facts.
Entities
Institutions
- arXiv