ARTFEED — Contemporary Art Intelligence

ConfBench: New Benchmark for Calibrating Vision-Language Models in Document Extraction

ai-technology · 2026-08-04

A team of researchers has launched ConfBench, the inaugural benchmark tailored for calibration in key information extraction (KIE) utilizing vision-language models (VLMs). This benchmark tackles the issue of reliable confidence scores in intelligent document processing (IDP), where models must determine whether to automate extraction or refer it for human analysis. Previous benchmarks have primarily featured clean, high-quality samples, resulting in insufficient low-accuracy areas for calibration evaluation. ConfBench was created by implementing 20 controlled degradation pipelines on a varied document collection, producing 1,346 variants and over 70,000 entity-level assessments across the entire accuracy range. The study assesses four proprietary and three open-weight VLMs using verbalized and log-probability confidence estimation methods across three input types. Key findings indicate: (i) the OCR+Image modality yields more precise confidence estimates; (ii) model capability is the primary influence, with confidence quality increasing consistently with capability within the Claude family. The research paper can be found on arXiv with the identifier 2608.01792.

Key facts

  • ConfBench is the first calibration-specific benchmark for key information extraction (KIE).
  • It was created by applying 20 controlled degradation pipelines to a diverse document set.
  • The benchmark includes 1,346 variants and over 70,000 entity-level evaluations.
  • Four proprietary and three open-weight vision-language models were evaluated.
  • Two confidence estimation methods were tested: verbalized and log-probability.
  • Three input modalities were compared.
  • OCR+Image modality leads to more accurate confidence estimates.
  • Model capability is the dominant factor in confidence quality, scaling monotonically within the Claude family.

Entities

Institutions

  • arXiv

Sources