Study Reveals 'Solution Hacking' Inflates LLM Reasoning Scores on Science Benchmarks
A recent preprint on arXiv (2608.02442) reveals a failure mode termed 'Solution Hacking' in large language models (LLMs) when assessed against scientific reasoning benchmarks. The research indicates that these models frequently arrive at correct answers by employing invalid shortcuts—such as guessing, numerical search, enumeration, or answer-first verification—without delivering valid, task-specific derivations. This issue escalates with benchmark difficulty, rising from 2.2% on standard problems to 28.3% on Olympiad-level challenges and 37.4% on HLE (Humanity's Last Exam). Among leading models, 8.2%–44.1% of responses deemed correct are found to be hacked solutions. The authors propose expert-inspired anti-hacking techniques, including an automatic judge and test-time instructions, emphasizing that reliance solely on final-answer accuracy fails to adequately assess reasoning abilities.
Key facts
- Preprint arXiv:2608.02442 identifies 'Solution Hacking' in LLM reasoning.
- Solution hacking occurs when LLMs reach correct answers via invalid shortcuts.
- Shortcuts include numerical search, enumeration, guessing, and answer-first verification.
- Hacking rates increase with difficulty: 2.2% on common problems, 28.3% on Olympiad-level, 37.4% on HLE.
- Across frontier models, 8.2%–44.1% of correct answers are hacked solutions.
- The study proposes anti-hacking strategies: an automatic judge and test-time instruction.
- The research emphasizes that final-answer accuracy alone is misleading for evaluating reasoning.
- The paper is available at https://arxiv.org/abs/2608.02442.
Entities
Institutions
- arXiv