Reachability vs. Realization: Auditing LLM Benchmark Gains
A new arXiv paper (2608.03219) challenges the assumption that benchmark gains reflect genuine LLM capability improvements. The authors argue that aggregate scores obscure whether models are reaching new answers or merely producing answers that were already within reach. They propose a question-level audit under fixed budgets, temperatures, and answer formats, distinguishing between 'realized' questions (correct under default deployment) and 'reachable' questions (correct via a specified probe within a budget). Testing inference-time layer routing across 43 model and task settings, they found that random routes match or exceed structured search under matched budgets, and that answer-blind procedures retain almost none of the gain, indicating the need for access to the correct answer. The study spans six cases from 0.5B to 31B parameter models, probing why reachable answers sometimes fail to appear. This work has significant implications for evaluating AI systems, suggesting that benchmark scores may overstate progress and that more granular auditing is necessary.
Key facts
- arXiv paper 2608.03219
- Question-level audit of LLM benchmarks
- Distinguishes 'realized' vs 'reachable' questions
- Random routes match or exceed structured search in 43 settings
- Answer-blind procedures retain almost none of the gain
- Requires access to correct answer for gains
- Spans six cases from 0.5B to 31B parameter models
- Challenges aggregate benchmark scores as evidence of capability
Entities
Institutions
- arXiv