LookBack: A New Method to Score LVLM Responses Using Visual Reference
A new research paper on arXiv (2608.11847v1) introduces LookBack, a training-free method for scoring responses of Large Vision-Language Models (LVLMs) by measuring how strongly each response token refers to image tokens. The authors argue that existing confidence-based metrics, adopted from LLMs, are insufficient for LVLMs because they primarily capture textual plausibility rather than agreement with the image. Their diagnostics show that removing the input image barely changes confidence-based selection. LookBack augments token likelihood with a visual lookback score, providing a lightweight measure of visual grounding. The method is evaluated across four benchmarks and three models, demonstrating improved scoring accuracy. The paper addresses the challenge of hallucinations in LVLMs, where models produce fluent responses ungrounded in the visual input.
Key facts
- arXiv:2608.11847v1
- Announce Type: cross
- Large Vision-Language Models (LVLMs) integrate visual perception with language generation
- LVLMs hallucinate against the image, producing responses ungrounded in what they see
- Existing confidence-based metrics from LLMs are insufficient for LVLMs
- Removing the input image barely changes confidence-based selection
- LookBack is a training-free LVLM response scoring method
- LookBack augments token likelihood with visual lookback score
- Evaluated across four benchmarks and three models
Entities
Institutions
- arXiv