ARTFEED — Contemporary Art Intelligence

ORCA-bench: Benchmarking LLM Agents for Oncall Root Cause Analysis

ai-technology · 2026-08-06

The newly introduced ORCA-bench serves to assess large language model (LLM) agents in production-level on-call situations. This benchmark integrates a live microservice system equipped with OpenTelemetry, utilizing six days' worth of metrics, logs, and traces accessible through genuine telemetry interfaces such as Prometheus, Jaeger, and OpenSearch via Grafana, along with complete source code. It features 1,079 root cause analysis (RCA) tasks that vary in report specificity, time-to-detection, and concurrent fault scenarios. Symptoms are curated and validated by expert SREs, while the evaluation of the LLM is independently re-assessed by humans, achieving a Cohen's kappa of 0.90. Among five leading agents, the highest RCA accuracy recorded is 25%.

Key facts

  • ORCA-bench is a benchmark for LLM agents in oncall root cause analysis.
  • It uses a live OpenTelemetry-instrumented microservice system.
  • The system exposes six days of metrics, logs, and traces via Prometheus, Jaeger, and OpenSearch through Grafana.
  • Full source-code access is provided.
  • There are 1,079 RCA tasks with variations in report specificity, time-to-detection, and co-occurring faults.
  • Ground-truth symptoms are curated by expert SREs.
  • LLM-as-judge is re-scored by humans with Cohen's kappa 0.90.
  • Best RCA Accuracy across five frontier agents is 25%.

Entities

Institutions

  • arXiv
  • OpenTelemetry
  • Prometheus
  • Jaeger
  • OpenSearch
  • Grafana

Sources