ARTFEED — Contemporary Art Intelligence

SciCode Benchmark Defects Underestimate LLM Scientific Coding Ability

ai-technology · 2026-08-06

A new study published on arXiv (2608.04975) reveals that the SciCode benchmark, a standard measure of language models' scientific-coding ability, contains defects that have led to an underestimation of their performance. The benchmark, which includes research-level problems requiring both scientific theory and numerical coding, is used in the Artificial Analysis Intelligence Index and government evaluations. Recent scores have plateaued, with top 2026 models achieving around 60% subproblem accuracy. An audit of all 65 test problems by domain experts found 263 defects, 192 of which (affecting 91% of main problems) cause correct solutions to be wrongly rejected due to non-reproducible gold answers, over-tight tolerances, or self-contradictory specifications. Notably, 78% of these score-suppressing defects require specialized physics or mathematics knowledge to identify. The findings suggest that language models' scientific-coding abilities are higher than previously reported, and the study calls for benchmark improvements.

Key facts

  • SciCode is a standard benchmark for scientific-coding ability of language models.
  • It is part of the Artificial Analysis Intelligence Index and used in government evaluations.
  • Recent scores plateaued around 60% subproblem accuracy for top 2026 models.
  • An audit of all 65 test problems found 263 defects.
  • 192 defects cause correct solutions to be wrongly rejected.
  • 91% of main problems are affected by these defects.
  • 78% of score-suppressing defects require specialized physics or math knowledge.
  • The study suggests language models' abilities are underestimated.

Entities

Institutions

  • arXiv
  • Artificial Analysis Intelligence Index

Sources