New Study: No Universal Signal Predicts Sample-Level LLM Regression Across Version Updates
An arXiv paper (2608.13607) explores the predictability of individual sample-level regression in large language models (LLMs) based on inference-time signals during model updates. The research contrasts single-model indicators (such as confidence, logit margin, and attention entropy) with cross-version indicators (including output KL divergence, likelihood drift, token-level KL, and representation drift) through a unified added-value test that assesses each signal's contribution beyond a confidence baseline. Analyzing six benchmarks across three task categories (multiple-choice question answering, math reasoning, and code generation) and six model update pairs, the findings reveal that signal effectiveness varies by task: confidence excels in MCQ and basic math, while likelihood/KL signals perform best in code generation and complex math. There is no one-size-fits-all signal for predicting regression across all tasks, underscoring the necessity for task-specific strategies. The paper has been submitted to arXiv.
Key facts
- Paper arXiv:2608.13607v1
- Announce Type: new
- Compares single-model and cross-version signals
- Unified added-value test isolates signal gains over confidence baseline
- Six benchmarks across three task families
- Six model update pairs
- Signal effectiveness is task-dependent
- Confidence strongest on MCQ and simpler math
- Likelihood/KL signals best for code generation and harder math
- No universal signal predicts regression
Entities
Institutions
- arXiv