ARTFEED — Contemporary Art Intelligence

Agent Leaderboards Rank Specialization, Not Capability: New Study

ai-technology · 2026-08-13

A recent investigation published on arXiv (2608.11323) indicates that leaderboards for agents utilized by enterprise professionals prioritize specialization over general skills. The study, which examined three benchmarks—TheAgentCompany, τ²-bench, and AppWorld—discovered that the main effect of agents contributes to less than 3% of the overall variance, while the interaction between agents and tasks contributes between 7% and 23%. This suggests that these leaderboards emphasize performance specific to tasks rather than overall ability. By employing Generalizability Theory variance decomposition alongside three estimators, the research points out the shortcomings of leaderboards, such as reliability challenges on complex tasks and adverse correlations between training and out-of-sample reliability. The paper is titled 'Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations.'

Key facts

  • Study published on arXiv (2608.11323) examines agent leaderboards.
  • Agent main effect accounts for less than 3% of total variance in all datasets.
  • Agent-by-task interaction accounts for 7-23% of variance.
  • Three benchmarks used: TheAgentCompany, τ²-bench, and AppWorld.
  • Four-facet Generalizability Theory variance decomposition employed.
  • Three estimators (Henderson Method-I, REML, Bayesian binomial GLMM) agree to three decimal places.
  • Aggregate reliability collapses on hardest task quartile: Eρ² on τ² action_checks falls from 0.752 to 0.000.
  • Training-cell reliability negatively correlates with held-out reliability (r = -0.90 on τ²).

Entities

Institutions

  • arXiv
  • TheAgentCompany
  • τ²-bench
  • AppWorld

Sources