DocMemo: Memory-Guided Retrieval for Multi-Modal Document Understanding
A new study has introduced DocMemo, a framework that seeks to improve how we discover evidence in long documents. This research, listed under arXiv paper number 2608.07067, addresses the limitations of existing systems that rely on static retrieval and inconsistent memory. DocMemo transforms the approach to reasoning through lengthy texts by focusing on dynamic evidence exploration. It incorporates a three-tier retrieval structure, which includes Document Schema Memory, Page Belief Memory, and Question Episodic Memory. These components are crafted to capture structural insights, evaluate relevance on the go, and outline reasoning paths tailored to specific queries. This framework aims to overcome the challenges faced by both single-round and iterative methods. The paper is categorized as a new announcement and is accessible online.
Key facts
- DocMemo is a memory-guided framework for multi-modal document understanding.
- It addresses limitations of static retrieval and fragile cross-round memory in long-document understanding.
- The framework maintains a tri-level retrieval state: Document Schema Memory, Page Belief Memory, and Question Episodic Memory.
- Document Schema Memory captures structural priors.
- Page Belief Memory provides dynamic relevance estimation.
- Question Episodic Memory tracks query-specific reasoning trajectories.
- The paper is available on arXiv with identifier 2608.07067.
- The announcement type is 'new'.
Entities
Institutions
- arXiv