BenchDrift: LLM Performance Varies with Problem Phrasing
A new study from arXiv (2608.11694) introduces BenchDrift, a tool that quantifies how rephrasing benchmark problems affects large language model (LLM) performance. The research shows that a single phrasing of a problem does not represent the entire space of possible phrasings, and that meaning-preserving variations can flip a model's answer in both directions—turning failures into successes and successes into failures. The study tests eight models across three benchmarks (GSM8K, MMLU, MATH-Hard) and finds that drift is substantial. Notably, phrasing sensitivity does not diminish as models improve; instead, it changes sign. Weaker models tend to benefit from rephrasing, while stronger models lose more than they gain. The tool generates variations along four axes: linguistic, referential, pragmatic, and structural. The findings highlight the fragility of benchmark scores and suggest that current evaluation methods may be misleading.
Key facts
- BenchDrift generates meaning-preserving variations of benchmark problems along four axes: linguistic, referential, pragmatic, and structural.
- The study tests eight models and three benchmarks: GSM8K, MMLU, MATH-Hard.
- Rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions.
- Phrasing sensitivity does not fade as models get better; instead, it changes sign.
- Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain.
- The paper is available on arXiv with ID 2608.11694.
- The announcement type is cross.
- The study demonstrates that benchmark scores are not robust to phrasing variations.
Entities
Institutions
- arXiv