ARTFEED — Contemporary Art Intelligence

LongRCA Bench: New Benchmark for Diagnosing Long-Horizon Agent Failures

ai-technology · 2026-08-18

A new benchmark called LongRCA Bench has been developed by researchers to identify failures in long-horizon agent executions. It includes 1,140 failed trajectories across five distinct domains, free from any injected errors, and offers independently evaluated human labels indicating the responsible role and the earliest critical root-cause step. The median trajectory consists of 145 steps, with the best baseline achieving merely 13.2% accuracy in pinpointing the exact root step, underscoring the challenge of this task. Additionally, the paper introduces Root-Cause Trajectory Attribution (RCTA), a method that operates without training, which identifies potential error steps from segment summaries and links them to prior handoffs. This research tackles the less-explored issue of failure attribution in long-horizon agents, where traditional outcome-level assessments fall short in tracing the source of errors. The benchmark can be accessed on arXiv (arXiv:2608.15242).

Key facts

  • LongRCA Bench comprises 1,140 failed trajectories across five domains.
  • No injected errors are present in the benchmark.
  • Human labels are provided for responsible role and earliest decisive root-cause step.
  • Median trajectory length is 145 steps.
  • Strongest baseline achieves only 13.2% exact root-step accuracy.
  • RCTA is a training-free method for root-cause attribution.
  • RCTA retrieves candidate error steps from segment summaries and traces them to earlier handoffs.
  • The paper is available on arXiv with ID 2608.15242.

Entities

Institutions

  • arXiv

Sources