ARTFEED — Contemporary Art Intelligence

TelemetrySuffBench: Benchmarking Failure-Origin Diagnosis in Agent Telemetry

ai-technology · 2026-08-11

A new benchmark called TelemetrySuffBench has been introduced to evaluate how well agent telemetry can pinpoint failure sources in systems with multiple components. It’s detailed in a paper on arXiv (2608.07899) and makes distinctions between detecting failures, locating faults, and safely stepping back when there’s insufficient evidence. The benchmark produces standard traces with delayed-binding faults and visualizes them using paired coarse views, seven-factor telemetry masks, and identical ambiguous origin pairs. Five cutting-edge language models were examined using consistent methods, including subgroup analyses and invalid-output checks. Results show that complete telemetry offers origin-step Top-1 accuracy ranging from 33.8% to 97.2%, while other views achieve high detection rates but limit origin accuracy to just 0.5%, highlighting significant gaps in robustness.

Key facts

  • TelemetrySuffBench is a controlled benchmark for failure-origin diagnosis in agent systems.
  • It separates failure detection, fault-origin localization, and safe abstention.
  • The benchmark constructs canonical multi-component traces with delayed-binding faults.
  • It uses paired coarse views, seven-factor telemetry masks, and exact-equal ambiguous origin pairs.
  • Five frontier language models were evaluated.
  • With full telemetry, origin-step Top-1 accuracy ranges from 33.8% to 97.2% across models.
  • Metadata, OpenTelemetry-compatible, and OpenInference-compatible views retain 99.5% to 100% detection F1.
  • Origin-step accuracy is limited to at most 0.5% with these views.
  • The paper is available on arXiv with ID 2608.07899.

Entities

Institutions

  • arXiv

Sources