ARTFEED — Contemporary Art Intelligence

New Benchmark SciFigBench Evaluates VLM Reliability on Scientific Figures

ai-technology · 2026-08-15

A novel diagnostic standard, SciFigBench, has been launched to assess vision-language models (VLMs) in the context of understanding scientific figures, with an emphasis on behavioral consistency amid uncertainty. This benchmark, outlined in arXiv paper 2608.13267, fills a void in current VLM evaluations that mainly focus on accuracy in perception and reasoning, often overlooking how models respond when visual data is either absent or misleading. SciFigBench features 250 figures, all meticulously annotated by humans, requiring over 600 hours of work across three evaluation criteria. The figures are enhanced through various methods, resulting in over 34,000 evaluation scenarios for rigorous testing. Additionally, the paper introduces the Admittance-Resistance-Inductance (A-R-I) framework to determine if models recognize inadequate information. This advancement is crucial for the AI and technology fields, offering a more thorough assessment approach for VLMs, which are increasingly utilized in scientific and educational applications.

Key facts

  • SciFigBench is a new diagnostic benchmark for VLM scientific figure understanding.
  • It evaluates perception, reasoning, and behavioral reliability under uncertainty.
  • The benchmark contains 250 figures with human annotations.
  • Over 600 hours of annotation effort were invested.
  • More than 34,000 evaluation setups were created for stress testing.
  • The A-R-I framework is proposed to evaluate model acknowledgment of insufficient information.
  • The paper is available on arXiv with ID 2608.13267.
  • Existing VLM benchmarks focus on perception and reasoning accuracy, not behavioral reliability.

Entities

Institutions

  • arXiv

Sources