ARTFEED — Contemporary Art Intelligence

New Study: No Universal Signal Predicts Sample-Level LLM Regression Across Version Updates

ai-technology · 2026-08-17

An arXiv paper (2608.13607) explores the predictability of individual sample-level regression in large language models (LLMs) based on inference-time signals during model updates. The research contrasts single-model indicators (such as confidence, logit margin, and attention entropy) with cross-version indicators (including output KL divergence, likelihood drift, token-level KL, and representation drift) through a unified added-value test that assesses each signal's contribution beyond a confidence baseline. Analyzing six benchmarks across three task categories (multiple-choice question answering, math reasoning, and code generation) and six model update pairs, the findings reveal that signal effectiveness varies by task: confidence excels in MCQ and basic math, while likelihood/KL signals perform best in code generation and complex math. There is no one-size-fits-all signal for predicting regression across all tasks, underscoring the necessity for task-specific strategies. The paper has been submitted to arXiv.

Key facts

  • Paper arXiv:2608.13607v1
  • Announce Type: new
  • Compares single-model and cross-version signals
  • Unified added-value test isolates signal gains over confidence baseline
  • Six benchmarks across three task families
  • Six model update pairs
  • Signal effectiveness is task-dependent
  • Confidence strongest on MCQ and simpler math
  • Likelihood/KL signals best for code generation and harder math
  • No universal signal predicts regression

Entities

Institutions

  • arXiv

Sources