Study Challenges Independence Assumption in Multi-Agent Reliability
A recent study published on arXiv examined the reliability of multi-agent systems through an analysis of 18,000 missions. It revealed that when two agents share tasks using identical models, there is a 90% likelihood that they will both fail if one does. By switching to a different model, this strong correlation was diminished. Furthermore, using models from different vendors showed no significant variance in performance, aligning with the researchers' pre-registered null hypothesis. These findings underscore the potential risks of assuming redundancy in systems reliant on the same model, highlighting important considerations for the reliability of AI systems.
Key facts
- Preprint on arXiv: 2608.12895
- Preregistered evaluation of 18,000 missions
- Two instances of one model co-fail on 90.0% of missions where either fails
- Log odds ratio 6.66, 95% CI [6.38, 7.00]
- Phi coefficient 0.916
- Substituting a different model reduces association in six of six contrasts
- Substituting a different vendor (model already different) does not reduce association
- Scored by deterministic code with no LLM judge
Entities
Institutions
- arXiv