ARTFEED — Contemporary Art Intelligence

In-Context Collapse in Vision-Language Models: A New Failure Mode and Mitigation

ai-technology · 2026-08-06

An arXiv paper (2608.02830) has uncovered an unexpected failure mode in vision-language models (VLMs). As the quantity of demonstrations in many-shot in-context learning (ICL) rises, certain VLMs suffer a significant and sometimes disastrous decline in accuracy, referred to as 'in-context collapse.' This issue is evident in synthetic classification, natural-image classification, and VQA benchmarks, with some models performing worse than chance while still producing coherent outputs. The research assesses an open VLM panel with parameters ranging from 0.5B to 11B and includes the advanced model Claude Sonnet 4.5, revealing a graded collapse. The study distinguishes between two abilities: resilience to increasing demonstrations and learning new rules in context, leading to three reproducible regimes. Through parameter-matched lesion-and-rescue experiments, the authors identify the vision-language integration pathway, specifically an adapter on the connector, as the source of the collapse. Proposed mitigation strategies aim to resolve this issue, challenging the prevalent belief that additional demonstrations always enhance ICL performance in VLMs and exposing a significant vulnerability in existing models.

Key facts

  • Many-shot in-context learning (ICL) in vision-language models (VLMs) can lead to a sharp accuracy drop as demonstrations accumulate.
  • This phenomenon, termed 'in-context collapse,' occurs in a subset of VLMs.
  • The collapse is observed across synthetic classification, natural-image classification, and VQA benchmarks.
  • Some models fall below chance accuracy while outputs remain well-formed.
  • The study includes an open VLM panel (0.5B–11B parameters) and Claude Sonnet 4.5.
  • Two capabilities are dissociable: robustness to accumulating demonstrations and learning a novel rule in context.
  • The collapse is causally localized to the vision-language integration pathway via an adapter on the connector.
  • The paper proposes mitigation strategies.
  • The paper is available on arXiv with ID 2608.02830.

Entities

Institutions

  • arXiv
  • Claude Sonnet 4.5

Sources