ARTFEED — Contemporary Art Intelligence

BenchDrift: LLM Performance Varies with Problem Phrasing

ai-technology · 2026-08-13

A new study from arXiv (2608.11694) introduces BenchDrift, a tool that quantifies how rephrasing benchmark problems affects large language model (LLM) performance. The research shows that a single phrasing of a problem does not represent the entire space of possible phrasings, and that meaning-preserving variations can flip a model's answer in both directions—turning failures into successes and successes into failures. The study tests eight models across three benchmarks (GSM8K, MMLU, MATH-Hard) and finds that drift is substantial. Notably, phrasing sensitivity does not diminish as models improve; instead, it changes sign. Weaker models tend to benefit from rephrasing, while stronger models lose more than they gain. The tool generates variations along four axes: linguistic, referential, pragmatic, and structural. The findings highlight the fragility of benchmark scores and suggest that current evaluation methods may be misleading.

Key facts

  • BenchDrift generates meaning-preserving variations of benchmark problems along four axes: linguistic, referential, pragmatic, and structural.
  • The study tests eight models and three benchmarks: GSM8K, MMLU, MATH-Hard.
  • Rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions.
  • Phrasing sensitivity does not fade as models get better; instead, it changes sign.
  • Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain.
  • The paper is available on arXiv with ID 2608.11694.
  • The announcement type is cross.
  • The study demonstrates that benchmark scores are not robust to phrasing variations.

Entities

Institutions

  • arXiv

Sources