ARTFEED — Contemporary Art Intelligence

Rubric-Based Grading: Smaller AI Models Match Costly Ones in Exam Scoring

ai-technology · 2026-08-19

Research indicates that smaller language models can assess open-ended exam responses with a reliability comparable to that of much pricier models, as long as they utilize a defined rubric for grading. This conclusion arises from the 'any-to-bench' design principle, where a leading model processes source documents once to extract questions and their rubrics. Subsequently, more affordable models manage all subsequent grading tasks. To validate this assertion, researchers examined six budget-friendly model configurations from two different families across three levels of reasoning effort. Each configuration responded to 24 open-ended exam questions and graded each answer three times, totaling 3,456 individual grades. The analysis demonstrated that the content of the answer predominantly influences the score, with answer identity accounting for 95.6% of score variance, while judge identity only explains 0.2%. A writer's reasoning effort can impact scores by up to 0.143 of full marks, and while a judge's reasoning effort also influences scores, the exact effect remains unspecified. This study underscores the potential for substantial cost reductions in educational assessments without compromising grading accuracy, as long as a clear rubric is utilized.

Key facts

  • Small language models grade open-ended exam answers as reliably as substantially more expensive models when grading against an explicit rubric.
  • The design principle is central to any-to-bench.
  • A frontier model reads source documents once at ingestion to extract each question and its rubric.
  • Lower-cost models perform all repeated grading work.
  • Six cost-efficient model configurations from two model families at three reasoning-effort levels were evaluated.
  • Each configuration answered 24 open-ended examination questions and graded every answer sheet three times, yielding 3,456 per-question grades.
  • Answer identity explains 95.6% of score variance; judge identity explains only 0.2%.
  • Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks.

Entities

Sources