ARTFEED — Contemporary Art Intelligence

CalibratedRubric: Adaptive Rubric Banks for LLM Evaluation

ai-technology · 2026-08-03

A new framework called CalibratedRubric has been developed by researchers to enhance the consistency of open-ended LLM evaluations through the selection of effective rubrics. This innovative approach integrates type-specific scoring, Bayesian filtering for rubric measurability, and the assembly of item response theory (IRT)-based banks. It assesses the measurability of each rubric using a Beta-Bernoulli agreement posterior and utilizes a submodular information-coverage objective for creating efficient rubric banks. Measurability filtering led to an increase in human-gold agreement on JudgmentBench from κ=0.604 to 0.743 across various benchmarks in finance, healthcare, general, and legal fields. Additionally, IRT-based greedy selection outperformed random selection in cross-fitted rank fidelity. The research can be found on arXiv with the identifier 2607.29252.

Key facts

  • CalibratedRubric is a task-adaptive framework for LLM evaluation.
  • It combines type-specific scoring, Bayesian rubric-measurability filtering, and IRT-based bank assembly.
  • Measurability filtering improved human-gold agreement on JudgmentBench from κ=0.604 to 0.743.
  • IRT-based greedy selection improved cross-fitted rank fidelity over random selection.
  • The framework was tested on financial, healthcare, general, and legal benchmarks.
  • The paper is available on arXiv with ID 2607.29252.
  • The approach addresses the cost and scalability of expert rubric curation.
  • It distinguishes measurable rubrics from informative ones.

Entities

Institutions

  • arXiv

Sources