OrchestraBench: New Benchmark Diagnoses Multi-Agent Orchestration Failures
A new standard, known as OrchestraBench, has been launched to assess failure modes, recovery processes, and the quality of decomposition in multi-agent orchestration systems. In contrast to conventional benchmarks that merely report task accuracy without identifying root problems, OrchestraBench employs a controlled, seed-reproducible failure-injection framework across standardized enterprise workflows. It features two key metrics: cascade radius and recovery per failure mode. The benchmark evaluates routing strategies through bootstrap confidence intervals and paired testing. In a diagnostic involving 26 gold-labelled cases, a keyword/flag router achieved 0% on adversarial scenarios with deceptive or absent surface flags, whereas an intent-reasoning model router attained a perfect score of 100%. The benchmark seeks to enhance understanding of multi-agent pipeline failures, facilitating improved routing and recovery strategies. The research paper can be found on arXiv with the identifier 2608.05263.
Key facts
- OrchestraBench is a new benchmark for multi-agent orchestration frameworks.
- It evaluates failure modes, recovery, and decomposition quality.
- It uses a controlled, seed-reproducible failure-injection harness.
- Primary metrics are cascade radius and per-failure-mode recovery.
- Routing policies are compared with bootstrap confidence intervals and paired tests.
- A keyword/flag router scored 0% on adversarial cases, while an intent-reasoning model router scored 100%.
- Controlled probes with a real Claude agent revealed three failure-handling tiers across five MAST modes.
- The paper is available on arXiv (2608.05263).
Entities
Institutions
- arXiv