ARTFEED — Contemporary Art Intelligence

Cheap Open-Weight LLMs Match Frontier Judges for Grading Math Proofs

ai-technology · 2026-08-04

A new study from arXiv (2608.00004) investigates whether cost-effective open-weight language models can reliably grade natural-language mathematical proofs, a task traditionally requiring expensive frontier models. The researchers evaluated three cheap models—GPT-OSS 120B, DeepSeek-V4 Flash, and Gemma-4 31B—on a 200-instance validation sample of IMO-GradingBench, comparing their pass/fail agreement with human graders against that of Claude Opus 4.7 and Gemini 3.1 Pro. Results showed that the cheap judges achieved statistically indistinguishable agreement rates from the frontier models, at up to 100 times lower cost. Surprisingly, a majority vote of the three cheap models did not improve over its strongest individual member, but requiring unanimous agreement (all-three-pass) yielded the highest pass-agreement and precision, with minimal run-to-run variance across four replicates. The findings suggest that budget-friendly open-weight models can serve as reliable judges for math proof evaluation, potentially reducing costs in AI evaluation pipelines. The study extends to the full 1000-instance benchmark, exploring consensus rules to optimize grading accuracy.

Key facts

  • Study on arXiv:2608.00004
  • Evaluates cheap open-weight LLMs as judges for natural-language math proofs
  • Models tested: GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B
  • Benchmark: IMO-GradingBench (200-instance validation, 1000-instance full)
  • Compared against Claude Opus 4.7 and Gemini 3.1 Pro
  • Cheap judges matched frontier models in agreement with human pass/fail decisions
  • Cost reduction up to 100x
  • Unanimous agreement (all-three-pass) achieved highest pass-agreement and precision

Entities

Institutions

  • arXiv
  • IMO-GradingBench

Sources