ARTFEED — Contemporary Art Intelligence

SafeIMG Benchmark Exposes AI Image Detection Gaps in Safety Scenarios

ai-technology · 2026-07-29

A new benchmark called SafeIMG, introduced in a preprint on arXiv, evaluates the ability of AI detectors to identify synthetic images in high-risk public- and individual-safety contexts. The benchmark spans 12 scenarios generated using GPT Image 2, including situations where fake visuals could threaten public safety or personal reputation. Unlike existing benchmarks that focus on generic imagery and image-level labels, SafeIMG provides human annotations localizing suspicious regions and explaining artifacts or physical inconsistencies. Tests on specialized synthetic-image detectors and vision-language models (VLMs) reveal significant shortcomings, highlighting the erosion of visual evidence authenticity in critical settings.

Key facts

  • SafeIMG is a safety-oriented benchmark for AI-generated image detection.
  • It covers 12 public- and individual-safety scenarios.
  • Images were generated using GPT Image 2.
  • Benchmark includes human annotations of suspicious regions and anomalies.
  • Evaluates both specialized detectors and vision-language models.
  • Existing detection benchmarks rarely examine safety-critical contexts.
  • Misleading visual content in safety scenarios carries substantial risks.
  • The study finds detectors often fail to recognize synthetic images in these contexts.

Entities

Institutions

  • arXiv

Sources