ARTFEED — Contemporary Art Intelligence

VLMs Rewrite Imperfect Text Instead of Transcribing Faithfully

ai-technology · 2026-07-27

A recent study indicates that Vision Language Models (VLMs) frequently transform flawed text into more credible versions instead of accurately transcribing it, a tendency that standard clean-text OCR benchmarks fail to detect. The researchers present FaithC4, a multilingual perturbation benchmark comprising 1,455 single-page documents in English, Chinese, and Korean, categorized into three types of perturbations: scramble, random substitution, and visually similar substitution. They assess 15 systems, including general-purpose VLMs, OCR-focused VLMs, and conventional OCR methods. General-purpose VLMs experience a degradation of up to 4.5 points in Word Error Rate (WER) under perturbation, while OCR-specialized VLMs decline by 0.2–2 points, and traditional OCR shows a drop of less than 0.6 points in English. Layer-by-layer analysis of Qwen3-VL-4B uncovers a consistent rewriting pattern. The paper can be found on arXiv (2607.21617).

Key facts

  • VLMs rewrite imperfect text into plausible forms instead of faithful transcription
  • FaithC4 benchmark includes 1,455 documents in English, Chinese, Korean
  • Three perturbation families: scramble, random substitution, visually similar substitution
  • 15 systems evaluated: general-purpose VLMs, OCR-specialized VLMs, traditional OCR
  • General-purpose VLMs degrade up to 4.5 WER points under perturbation
  • OCR-specialized VLMs degrade 0.2–2 WER points
  • Traditional OCR degrades less than 0.6 WER points on English
  • Layer-by-layer probing of Qwen3-VL-4B shows consistent rewriting behavior

Entities

Institutions

  • arXiv

Sources