Study Attributes Failure Modes in Multi-Page Visually Rich Document Understanding
An empirical investigation published on arXiv (submission 2608.07943) delves into the shortcomings of multi-page visually rich document understanding (MP-VRDU) systems, identifying three primary failure modes: representation, selection, and reasoning. The researchers examined each mode by manipulating one while keeping the others constant, utilizing a dataset focused on multi-page document comprehension. Notable results reveal that while vision is essential, it cannot substitute for text extraction; accuracy is limited by missing pages, and the presence of distractors has minimal impact. Additionally, reasoners struggle to synthesize evidence from multiple pages, even when all information is available. The study offers recommendations for developing these systems within a fixed computational budget and falls under the category of Computer Science > Artificial Intelligence on arXiv.
Key facts
- The study is titled 'Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution'.
- It was published on arXiv with ID 2608.07943.
- The research identifies three failure modes: representation, selection, and reasoning.
- The method involves intervening on one failure mode while holding others fixed.
- Findings show vision is necessary but does not replace text extraction.
- Missing pages bound accuracy while distractors cost little.
- Reasoners fail to integrate evidence across pages even when fully supplied.
- Prompting can shift reasoning behavior substantially, with trade-offs.
Entities
Institutions
- arXiv