ARTFEED — Contemporary Art Intelligence

Item Response Theory Applied to AI Safety Benchmarks

ai-technology · 2026-06-23

A recent paper on arXiv (2608.05086) utilizes Item Response Theory (IRT), a statistical approach from psychometrics, to examine safety standards for large language models. This research applies IRT models to eight safety benchmarks involving 192 language models, representing the most extensive psychometric assessment of LLM safety evaluations to date. The authors uncover three key factors—refusal strictness, truthfulness, and contextual harm—that account for the majority of variance among models across benchmarks. Additionally, they demonstrate that items chosen through psychometric methods yield full benchmark scores with less error compared to randomly selected subsets of the same size, with approximately ten adaptive items being adequate. The study also tackles issues related to benchmark duplication, high correlations, and potential model sandbagging during evaluations, aiming to enhance the reliability and clarity of safety assessments.

Key facts

  • The paper is titled 'Item Response Theory for AI Safety' and is available on arXiv (2608.05086).
  • IRT models were fitted to eight safety benchmarks across 192 language models.
  • Three factors: refusal strictness, truthfulness, and contextual harm explain most variance.
  • Psychometrically selected items recover full benchmark scores with lower error than random subsets.
  • Roughly ten adaptive items are sufficient for accurate measurement.
  • The study addresses benchmark duplication, high correlation, and sandbagging.
  • This is the largest psychometric analysis of LLM safety evaluations to date.

Entities

Institutions

  • arXiv

Sources