LLM Rankings Shift with Token Budgets: Study Finds Non-Monotone Accuracy and Model Complementarity
A recent investigation disputes the belief that rankings of large language models (LLMs) are consistent under varying inference conditions. This study, available on arXiv (2608.12150), methodically alters the token generation budget—from 64 to 4,096 tokens—across seven levels. Four models were assessed using three reasoning benchmarks, resulting in 56,476 inferences. The results indicate that 3–19% of items demonstrate non-monotonic behavior, with accuracy declining as the budget increases, despite controlling for truncation. This behavior is specific to individual models, showing 6% to 14% overlap across models. Rankings shift across budgets on all benchmarks, achieving statistical significance (p < 0.01, McNemar). An oracle analysis reveals model complementarity of up to +27.8 percentage points, particularly at lower budgets. A budget-aware router addresses 14.1% of the oracle gap across domains, while budget features improve within-domain performance (+1.6 to +5.7 percentage points) but negatively impact transfer (-1.2 percentage points). These findings advocate for evaluation protocols that account for budgets and suggest that inference budgets should be a factor in model selection. This research was conducted by a team of researchers and is newly published on arXiv.
Key facts
- Study posted on arXiv (2608.12150) challenges stable LLM rankings.
- Varies token generation budget across seven levels (64–4,096).
- Evaluates four models on three reasoning benchmarks (56,476 inferences).
- 3–19% of items show non-monotone accuracy (decrease with more budget).
- Non-monotone behavior is model-specific (cross-model overlap 6–14%).
- Model rankings reverse across budgets on all benchmarks (p < 0.01, McNemar).
- Oracle analysis shows model complementarity up to +27.8 percentage points.
- Budget-aware router captures 14.1% of oracle gap cross-domain.
- Budget features help within-domain (+1.6 to +5.7 pp) but hurt transfer (-1.2 pp).
Entities
Institutions
- arXiv