Shortcut Cascades in Clinical Multi-Agent Systems: Benchmark Gaming Risks
A recent study published on arXiv (2608.03744) examines the susceptibility of language-model agent committees, utilized for clinical decision support, to shortcuts that benchmarks may reward but clinicians would disregard. The analysis includes seven cohorts from six public datasets, covering text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert), and tabular ICU records (SUPPORT2). While Gemini committees show resistance to these cues individually, a socially plausible shortcut emerges: when two peers provide the same incorrect answer, the holdout adopts it 38% of the time, as does a misleading 'pre-screen' flag across both capability levels. The study underscores the risks of cascading errors in multi-agent systems and emphasizes the necessity for effective oversight mechanisms.
Key facts
- The study is published on arXiv with ID 2608.03744.
- It examines clinical multi-agent systems for decision support.
- Seven cohorts across six public datasets were used.
- Datasets include MedQA-USMLE, MedMCQA, MIMIC-CXR reports, NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert, and SUPPORT2.
- Gemini committees resist shortcuts in isolation with flip rates of 5-16%.
- Socially plausible shortcut: when two peers assert the same wrong answer, the holdout adopts it in 38% of cases.
- A false 'pre-screen' system flag also leads to adoption in 38% of cases.
- Oversight agents: a gate has 100% false-positive rate; a same-lineage judge achieves precision 100% and recall 93% on text but fails in imaging.
Entities
Institutions
- arXiv