Causal Visual Memory Audit Reveals Flaws in Multimodal AI Recall
A new study from arXiv introduces the Causal Visual Memory Audit (CVMA), a framework testing whether stateful multimodal assistants can safely forget visual information across dialog turns. The research finds that current attention-based visual-KV eviction methods often rank future-useful visual regions worse than random, despite significant selection headroom. On VisDial and ConvBench benchmarks, aggregate scores mask this failure when later turns do not require vision. Controlled and stock-generated histories reveal that assistant-text KV can substitute for image KV for already stated facts, but not reliably for unstated ones. The study challenges assumptions in attention-guided eviction and highlights risks in long-term visual memory for AI assistants.
Key facts
- CVMA is a paired single-prefill framework for auditing visual memory in multimodal assistants.
- Current attention-based visual-KV eviction ranks future-useful regions worse than random.
- Diagnostic marginal-utility control shows substantial selection headroom.
- Tests conducted on VisDial and ConvBench benchmarks.
- Aggregate scores hide failure when later turns do not need vision.
- Assistant-text KV can replace image KV for stated facts but not unstated ones.
- The study is published on arXiv with ID 2607.25467.
- The research questions the safety of forgetting visual evidence in long dialog turns.
Entities
Institutions
- arXiv