ARTFEED — Contemporary Art Intelligence

Shortcut Cascades in Clinical Multi-Agent Systems: Benchmark Gaming Risks

ai-technology · 2026-08-06

A recent study published on arXiv (2608.03744) examines the susceptibility of language-model agent committees, utilized for clinical decision support, to shortcuts that benchmarks may reward but clinicians would disregard. The analysis includes seven cohorts from six public datasets, covering text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert), and tabular ICU records (SUPPORT2). While Gemini committees show resistance to these cues individually, a socially plausible shortcut emerges: when two peers provide the same incorrect answer, the holdout adopts it 38% of the time, as does a misleading 'pre-screen' flag across both capability levels. The study underscores the risks of cascading errors in multi-agent systems and emphasizes the necessity for effective oversight mechanisms.

Key facts

  • The study is published on arXiv with ID 2608.03744.
  • It examines clinical multi-agent systems for decision support.
  • Seven cohorts across six public datasets were used.
  • Datasets include MedQA-USMLE, MedMCQA, MIMIC-CXR reports, NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert, and SUPPORT2.
  • Gemini committees resist shortcuts in isolation with flip rates of 5-16%.
  • Socially plausible shortcut: when two peers assert the same wrong answer, the holdout adopts it in 38% of cases.
  • A false 'pre-screen' system flag also leads to adoption in 38% of cases.
  • Oversight agents: a gate has 100% false-positive rate; a same-lineage judge achieves precision 100% and recall 93% on text but fails in imaging.

Entities

Institutions

  • arXiv

Sources