Stable Miscalibration in LLMs: High-Confidence Errors Persist Under Perturbations
A recent paper on arXiv (ID 2608.13591) explores the origins of high-confidence errors in large language models, suggesting these issues may arise from stable miscalibration instead of unstable internal inference. The researchers utilize two diagnostic methods: a label-aware output-level audit score that assesses confidence variation across domains and identifies overconfident errors against a forced-answer baseline, along with an internal sensitivity probe that evaluates hidden-state shifts. In a multi-domain binary factual audit set, the audit score indicates where self-critique that accounts for abstention minimizes decision loss, although direct labeled baselines show a stronger ranking of the same improvement. Self-critical prompting effectively lowers hidden-state sensitivity across layers in three open-weight models, supporting prompt-induced local stabilization without confirming calibration.
Key facts
- Paper ID: arXiv:2608.13591
- Announcement type: new
- Focus: stable miscalibration in large language models
- Two diagnostics: output-level audit score and internal sensitivity probe
- Audit score ranks domains by confidence variation and overconfident mistakes
- Self-critical prompting reduces hidden-state sensitivity across layers in three open-weight models
- Findings support prompt-induced local stabilization, not purely output-level abstention
- Does not imply calibration
Entities
Institutions
- arXiv