Majority Vote Reduces Accuracy on Most GPQA Diamond Problems for Small LLMs
A new preprint on arXiv (2608.11403) reveals that using majority voting, a common inference method, actually harms the performance of smaller instruction-tuned language models on tough scientific tasks. In the GPQA Diamond benchmark, which includes 198 graduate-level science questions, Qwen2.5-7B saw a 56.6% drop in accuracy, while Llama-3-8B experienced a 65.7% decline. Qwen was the primary focus, with Llama confirming similar results from a near-chance baseline. The study validated all four hypotheses, based on 151 confirmatory problems and insights from 47 exploratory cases. Moreover, a theoretical upper limit showed Qwen and Llama could achieve 14 and 17 accuracy points more than N = 1, respectively, but no method without a verifier reached this limit, highlighting the need for advanced inference methods for smaller models dealing with complex reasoning.
Key facts
- Majority voting reduces per-problem accuracy on a majority of GPQA Diamond problems for Qwen2.5-7B (56.6%) and Llama-3-8B (65.7%).
- The effect was pre-registered on a 151-problem confirmatory split after being observed on 47 exploratory problems.
- All four confirmatory hypotheses passed.
- A grid oracle routing each problem to the best N across {1, 2, 4, 8, 16, 32, 64} marks a theoretical upper bound 14 accuracy points above N = 1 for Qwen and 17 for Llama.
- The oracle bound requires ground truth and is not a deployable method.
- No verifier-free gate reaches the oracle bound.
- The study uses the GPQA Diamond benchmark of 198 graduate-level science questions.
- The paper is available on arXiv with identifier 2608.11403.
Entities
Institutions
- arXiv