ARTFEED — Contemporary Art Intelligence

Verbalized Uncertainty for Risk-Controlled Deferral in Small Language Models

other · 2026-08-06

A recent preprint on arXiv (2608.05064) explores the role of expressed confidence in facilitating risk-controlled deferral for small open-weight language models, which are being utilized more frequently in private, offline, and budget-sensitive environments. The research assesses eleven instruction-tuned models from three different categories, with parameter sizes ranging from 0.5B to 14B, across ARC-Challenge and TruthfulQA, resulting in 25,168 local predictions. The authors present three key theoretical findings: strict monotone calibration maintains the risk-coverage frontier and error-detection AUROC; temperature scaling is ineffective for models with confidence exceeding one half while accuracy declines; and a Clopper-Pearson method transforms a 200-question calibration dataset into a finite-sample risk certificate under the i.i.d. assumption. Notably, eight out of 22 model-task combinations reached the temperature-scaling infeasibility limit within the study's parameters. The paper can be accessed on arXiv with the identifier 2608.05064.

Key facts

  • The paper is arXiv:2608.05064, announced as a cross-type submission.
  • It studies verbalized confidence for risk-controlled deferral in small language models.
  • Eleven instruction-tuned models from three families were evaluated, ranging from 0.5B to 14B parameters.
  • Evaluations were conducted on ARC-Challenge and TruthfulQA.
  • A total of 25,168 local predictions were generated.
  • Three theoretical results are presented: monotone calibration preserves risk-coverage frontier and AUROC; temperature scaling cannot calibrate certain models; Clopper-Pearson procedure provides finite-sample risk certificates.
  • Empirically, eight of 22 model-task pairs hit the temperature-scaling infeasibility floor.
  • The paper is available at https://arxiv.org/abs/2608.05064.

Entities

Institutions

  • arXiv

Sources