Reasoning Models Fail to Allocate Test-Time Compute Across Questions
A new arXiv paper (2608.07968) introduces an exam-style evaluation framework to study how reasoning language models allocate a shared token budget across multiple questions with varying difficulty and point values. The study finds that models, including several open and frontier reasoning models, fail to strategically distribute compute, instead behaving as greedy sequential solvers that prioritize by presentation order and front-load effort on early questions, remaining insensitive to value. These tendencies become more pronounced as the number of questions grows. The research highlights a significant limitation in current reasoning models' ability to manage end-to-end cost or latency constraints when solving multiple problems.
Key facts
- Paper arXiv:2608.07968 introduces an exam-style evaluation framework.
- Models must distribute one shared token budget across questions with different difficulty and point values.
- Several open and frontier reasoning models were evaluated.
- Models fail to allocate a shared budget strategically.
- Models behave as greedy sequential solvers, prioritizing by presentation order.
- Models front-load effort on early questions and are insensitive to value.
- Tendencies become more pronounced as the number of questions grows.
- The study addresses test-time compute allocation across questions, not just one at a time.
Entities
Institutions
- arXiv