VLMs Rewrite Imperfect Text Instead of Transcribing Faithfully
A recent study indicates that Vision Language Models (VLMs) frequently transform flawed text into more credible versions instead of accurately transcribing it, a tendency that standard clean-text OCR benchmarks fail to detect. The researchers present FaithC4, a multilingual perturbation benchmark comprising 1,455 single-page documents in English, Chinese, and Korean, categorized into three types of perturbations: scramble, random substitution, and visually similar substitution. They assess 15 systems, including general-purpose VLMs, OCR-focused VLMs, and conventional OCR methods. General-purpose VLMs experience a degradation of up to 4.5 points in Word Error Rate (WER) under perturbation, while OCR-specialized VLMs decline by 0.2–2 points, and traditional OCR shows a drop of less than 0.6 points in English. Layer-by-layer analysis of Qwen3-VL-4B uncovers a consistent rewriting pattern. The paper can be found on arXiv (2607.21617).
Key facts
- VLMs rewrite imperfect text into plausible forms instead of faithful transcription
- FaithC4 benchmark includes 1,455 documents in English, Chinese, Korean
- Three perturbation families: scramble, random substitution, visually similar substitution
- 15 systems evaluated: general-purpose VLMs, OCR-specialized VLMs, traditional OCR
- General-purpose VLMs degrade up to 4.5 WER points under perturbation
- OCR-specialized VLMs degrade 0.2–2 WER points
- Traditional OCR degrades less than 0.6 WER points on English
- Layer-by-layer probing of Qwen3-VL-4B shows consistent rewriting behavior
Entities
Institutions
- arXiv