ARTFEED — Contemporary Art Intelligence

OrchestraBench: New Benchmark Diagnoses Multi-Agent Orchestration Failures

ai-technology · 2026-08-07

A new standard, known as OrchestraBench, has been launched to assess failure modes, recovery processes, and the quality of decomposition in multi-agent orchestration systems. In contrast to conventional benchmarks that merely report task accuracy without identifying root problems, OrchestraBench employs a controlled, seed-reproducible failure-injection framework across standardized enterprise workflows. It features two key metrics: cascade radius and recovery per failure mode. The benchmark evaluates routing strategies through bootstrap confidence intervals and paired testing. In a diagnostic involving 26 gold-labelled cases, a keyword/flag router achieved 0% on adversarial scenarios with deceptive or absent surface flags, whereas an intent-reasoning model router attained a perfect score of 100%. The benchmark seeks to enhance understanding of multi-agent pipeline failures, facilitating improved routing and recovery strategies. The research paper can be found on arXiv with the identifier 2608.05263.

Key facts

  • OrchestraBench is a new benchmark for multi-agent orchestration frameworks.
  • It evaluates failure modes, recovery, and decomposition quality.
  • It uses a controlled, seed-reproducible failure-injection harness.
  • Primary metrics are cascade radius and per-failure-mode recovery.
  • Routing policies are compared with bootstrap confidence intervals and paired tests.
  • A keyword/flag router scored 0% on adversarial cases, while an intent-reasoning model router scored 100%.
  • Controlled probes with a real Claude agent revealed three failure-handling tiers across five MAST modes.
  • The paper is available on arXiv (2608.05263).

Entities

Institutions

  • arXiv

Sources