ARTFEED — Contemporary Art Intelligence

DocTrace: Hierarchical Evidence Graph Reasoning for LongDocVQA

ai-technology · 2026-08-06

A new approach called DocTrace has been developed to improve how we trace information in Long Document Visual Question Answering (LongDocVQA). This method, detailed in an arXiv paper (2608.03292), shifts the focus to reasoning through an explicit evidence graph instead of just guessing answers. DocTrace features a structured process that includes evidence localization, document parsing, and reasoning to clarify where the evidence comes from. It employs a two-step training process, beginning with joint Supervised Fine-Tuning (SFT). This framework aims to overcome limitations seen in current techniques, like end-to-end Multimodal Large Language Models (MLLMs) and retrieval-augmented generation (RAG) systems, by ensuring that evidence is accurately represented and verified, ultimately enhancing accuracy in LongDocVQA tasks.

Key facts

  • DocTrace is a hierarchical framework for LongDocVQA.
  • It casts LongDocVQA as an explicit evidence graph reasoning problem.
  • The framework performs evidence localization, structured document parsing, and evidence graph reasoning.
  • It enables explicit evidence provenance.
  • Training uses a two-stage framework starting with joint Supervised Fine-Tuning (SFT).
  • The paper is available on arXiv with ID 2608.03292.
  • Existing approaches include end-to-end MLLMs, RAG pipelines, and document agents.
  • The goal is to improve answer accuracy and traceability.

Entities

Institutions

  • arXiv

Sources