ARTFEED — Contemporary Art Intelligence

LLM Reliability Varies Significantly Under Meaning-Preserving Paraphrases

ai-technology · 2026-07-29

A recent study published on arXiv (2607.22554) examines the reactions of large language models (LLMs) to paraphrased questions that maintain their meaning. Researchers analyzed 13 models across four benchmarks and discovered that, although the overall accuracy remains relatively stable, the behavior at the instance level is inconsistent. Models often switch between right and wrong responses based on how questions are phrased, with mismatch rates surpassing 23%. Furthermore, when focusing on initially correct answers, the failure rates are even more pronounced, suggesting that the accuracy of a single prompt is not a reliable measure of performance.

Key facts

  • Study examines LLM behavior under meaning-preserving paraphrases.
  • Four benchmarks and 13 models were tested.
  • Instance-level answer mismatch rates exceed 23%.
  • Single-prompt correctness is a poor reliability indicator.
  • Overall accuracy changes modestly across paraphrases.
  • Answer flip rates are higher for originally correct questions.
  • Tasks include factual QA and mathematical reasoning.
  • Model outputs depend on exact wording of prompt.

Entities

Sources