ARTFEED — Contemporary Art Intelligence

Aggregate Benchmark Scores Hide Item-Level Regressions in GPT API Migrations

ai-technology · 2026-08-19

A recent paper on arXiv explores how aggregate benchmark scores can mask regressions at the item level during software migrations between versions of large language model APIs. The research focuses on three upgrades from GPT-5.4 to GPT-5.6, analyzing 900 benchmark items across diverse tasks by conducting 50 queries per item for each model. Items were categorized as improved, regressed, equivalent, or inconclusive through statistical analysis. Findings reveal that while overall migrations may appear successful, they can conceal significant item-level declines. The study highlights the necessity of item-level evaluations to grasp model changes and is available as preprint 2608.17719 on arXiv, pending peer review.

Key facts

  • The paper critiques reliance on aggregate benchmark scores for LLM API migration decisions.
  • It studies the GPT-5.4 to GPT-5.6 Sol product sequence with three pairwise upgrades.
  • Researchers queried 900 public benchmark items, each 50 times per model.
  • Benchmark categories included graduate-level knowledge, olympiad mathematics, and instruction following.
  • Items were classified as reliably improved, reliably regressed, practically equivalent, or inconclusive.
  • The classification used false-discovery-rate control and a practical-significance threshold.
  • Results were calibrated against a label-permutation null.
  • Reliable improvements and reliable regressions coexisted across all nine migration-benchmark cells.

Entities

Sources