ARTFEED — Contemporary Art Intelligence

Study Finds Shorter Reasoning Can Maintain Accuracy in LLMs

ai-technology · 2026-08-06

A recent investigation published on arXiv (2608.03401) examines the impact of shortening reasoning duration in large language models (LLMs) on their accuracy. Researchers conducted paired tests on the same questions at equivalent reasoning intervals, utilizing 198 GPQA Diamond and 500 MMLU-Pro queries. They implemented a numeric/concision prompt that sets a token limit for Qwen3-14B and the trained configurations of gpt-oss-20b and -120b. The Qwen prompt led to a 12-17% reduction in reasoning traces, with minimal and varied accuracy shifts at the same token limits. Notably, a concise/early-answer directive improved MMLU-Pro accuracy by 3.8 percentage points at 512 tokens. The study indicates that shorter reasoning may enhance accuracy, challenging the notion that longer reasoning is always superior.

Key facts

  • Study evaluates reasoning interfaces for LLMs
  • Uses 198 GPQA Diamond and 500 MMLU-Pro questions
  • Tests numeric/concision prompt on Qwen3-14B
  • Tests trained effort settings of gpt-oss-20b and -120b
  • Qwen prompt shortens reasoning traces by 12-17%
  • Accuracy changes at matched token limits are small and mixed
  • Concise/early-answer instruction raises MMLU-Pro accuracy by 3.8 points at 512 tokens
  • gpt-oss low/medium effort answers are 14.5-26.3 points more accurate than high effort

Entities

Institutions

  • arXiv

Sources