ARTFEED — Contemporary Art Intelligence

Causal Visual Memory Audit Reveals Flaws in Multimodal AI Recall

ai-technology · 2026-07-29

A new study from arXiv introduces the Causal Visual Memory Audit (CVMA), a framework testing whether stateful multimodal assistants can safely forget visual information across dialog turns. The research finds that current attention-based visual-KV eviction methods often rank future-useful visual regions worse than random, despite significant selection headroom. On VisDial and ConvBench benchmarks, aggregate scores mask this failure when later turns do not require vision. Controlled and stock-generated histories reveal that assistant-text KV can substitute for image KV for already stated facts, but not reliably for unstated ones. The study challenges assumptions in attention-guided eviction and highlights risks in long-term visual memory for AI assistants.

Key facts

  • CVMA is a paired single-prefill framework for auditing visual memory in multimodal assistants.
  • Current attention-based visual-KV eviction ranks future-useful regions worse than random.
  • Diagnostic marginal-utility control shows substantial selection headroom.
  • Tests conducted on VisDial and ConvBench benchmarks.
  • Aggregate scores hide failure when later turns do not need vision.
  • Assistant-text KV can replace image KV for stated facts but not unstated ones.
  • The study is published on arXiv with ID 2607.25467.
  • The research questions the safety of forgetting visual evidence in long dialog turns.

Entities

Institutions

  • arXiv

Sources