ARTFEED — Contemporary Art Intelligence

Forensic Audit Reveals Reproducibility Failures in Radiology AI Benchmark

ai-technology · 2026-07-30

A recent audit focused on a chest-radiograph vision-language model (VLM) benchmark found significant discrepancies between the model's released version and its original guidelines. This study, shared on arXiv, looked into various aspects like prompt bindings, DICOM metadata, and the completeness of outputs. It also assessed the consistency of releases across different datasets such as provider APIs and statistical codes. Out of 300 planned model-prompt interactions, 297 yielded usable reports. The review involved 30 studies with 28 patients, revealing that four MONOCHROME1 images did not have the required polarity inversion. Additionally, an unverified extractor restricted five reports to 4000 characters. The audit underlined the critical need for rigorous reproducibility checks in AI benchmarks for medical imaging.

Key facts

  • Retrospective forensic reproducibility audit of a chest-radiograph VLM benchmark
  • 300 planned model-prompt calls; 297 yielded nonempty reports
  • 60 Claude calls labeled A/B were executed with the same C prompt
  • 30 studies represented 28 patients
  • 4 MONOCHROME1 images rendered without required polarity inversion
  • Dataset split membership not retained
  • Unvalidated extractor truncated 5 reports to 4000 characters
  • No model was called again; no image or report was newly annotated

Entities

Institutions

  • arXiv

Sources