TelemetrySuffBench: Benchmarking Failure-Origin Diagnosis in Agent Telemetry
A new benchmark called TelemetrySuffBench has been introduced to evaluate how well agent telemetry can pinpoint failure sources in systems with multiple components. It’s detailed in a paper on arXiv (2608.07899) and makes distinctions between detecting failures, locating faults, and safely stepping back when there’s insufficient evidence. The benchmark produces standard traces with delayed-binding faults and visualizes them using paired coarse views, seven-factor telemetry masks, and identical ambiguous origin pairs. Five cutting-edge language models were examined using consistent methods, including subgroup analyses and invalid-output checks. Results show that complete telemetry offers origin-step Top-1 accuracy ranging from 33.8% to 97.2%, while other views achieve high detection rates but limit origin accuracy to just 0.5%, highlighting significant gaps in robustness.
Key facts
- TelemetrySuffBench is a controlled benchmark for failure-origin diagnosis in agent systems.
- It separates failure detection, fault-origin localization, and safe abstention.
- The benchmark constructs canonical multi-component traces with delayed-binding faults.
- It uses paired coarse views, seven-factor telemetry masks, and exact-equal ambiguous origin pairs.
- Five frontier language models were evaluated.
- With full telemetry, origin-step Top-1 accuracy ranges from 33.8% to 97.2% across models.
- Metadata, OpenTelemetry-compatible, and OpenInference-compatible views retain 99.5% to 100% detection F1.
- Origin-step accuracy is limited to at most 0.5% with these views.
- The paper is available on arXiv with ID 2608.07899.
Entities
Institutions
- arXiv