DICA: New Method Reduces Hallucinations in Multimodal LLMs
A study proposes Dual-Indicator Guided Contrastive Alignment (DICA) to reduce hallucinations in multimodal large language models (MLLMs). The method tracks Visual Attention Entropy (VAE) and Output Image Correlation (OIC) during inference. Abnormal VAE increases or OIC decreases trigger targeted contrastive alignment to restore visual grounding. Experiments across multiple benchmarks show consistent improvements.
Key facts
- DICA uses two indicators: VAE and OIC.
- VAE reflects concentration of visual attention.
- OIC measures dependence of outputs on visual input.
- Abnormal VAE increase or OIC decrease triggers alignment.
- Method aims to mitigate attention drift and underutilization of visual evidence.
- Experiments conducted across multiple benchmarks.
- Human visual reasoning follows coarse-to-fine attention process.
- MLLMs may deviate from this pattern.
Entities
—