Logit Lens Reveals Why LVLMs Hallucinate Objects Despite Strong Visual Attention
A recent paper on arXiv (2608.07302) disputes the common belief that inadequate visual attention is the root cause of object hallucination in Large Vision-Language Models (LVLMs). The researchers discovered that both genuine and hallucinated objects receive comparable levels of visual attention in the model's mid-to-late layers, indicating that the problem lies not in the quantity of attention but in its focus and reasoning. Utilizing Logit Lens to decode visual features in high-attention areas, they found that real object regions could be accurately mapped to target object tokens, unlike those of hallucinated objects. The study uncovers two separate hallucination mechanisms: visual uncertainty, which arises from semantically similar regions and can be addressed by masking, and contextual prior, influenced by strong co-occurrence patterns. The authors suggest strategies to detect and reduce these hallucinations, providing a fresh outlook on enhancing LVLM reliability, which is crucial for advancing AI systems in visual comprehension.
Key facts
- Paper arXiv:2608.07302, announced as cross type.
- LVLMs often hallucinate objects absent from the image.
- Both real and hallucinated objects receive equally strong visual attention in mid-to-late layers.
- Logit Lens decoding reveals real objects decode to target tokens, hallucinated ones do not.
- Two hallucination mechanisms identified: visual uncertainty and contextual prior.
- Masking semantically similar regions eliminates hallucination from visual uncertainty.
- Contextual prior is triggered by strong co-occurrence priors.
- Proposed methods detect and mitigate object hallucination in LVLMs.
Entities
Institutions
- arXiv