CalibratedRubric: Adaptive Rubric Banks for LLM Evaluation
A new framework called CalibratedRubric has been developed by researchers to enhance the consistency of open-ended LLM evaluations through the selection of effective rubrics. This innovative approach integrates type-specific scoring, Bayesian filtering for rubric measurability, and the assembly of item response theory (IRT)-based banks. It assesses the measurability of each rubric using a Beta-Bernoulli agreement posterior and utilizes a submodular information-coverage objective for creating efficient rubric banks. Measurability filtering led to an increase in human-gold agreement on JudgmentBench from κ=0.604 to 0.743 across various benchmarks in finance, healthcare, general, and legal fields. Additionally, IRT-based greedy selection outperformed random selection in cross-fitted rank fidelity. The research can be found on arXiv with the identifier 2607.29252.
Key facts
- CalibratedRubric is a task-adaptive framework for LLM evaluation.
- It combines type-specific scoring, Bayesian rubric-measurability filtering, and IRT-based bank assembly.
- Measurability filtering improved human-gold agreement on JudgmentBench from κ=0.604 to 0.743.
- IRT-based greedy selection improved cross-fitted rank fidelity over random selection.
- The framework was tested on financial, healthcare, general, and legal benchmarks.
- The paper is available on arXiv with ID 2607.29252.
- The approach addresses the cost and scalability of expert rubric curation.
- It distinguishes measurable rubrics from informative ones.
Entities
Institutions
- arXiv