DrIG: A New Generative Framework for Universal Multimodal Retrieval
A new research paper on arXiv (2608.12987) introduces DrIG, a generative framework for universal multimodal retrieval that uses dual-role identifiers to support diverse retrieval tasks across text, image, and mixed image-text items. The paper, announced as a cross-type submission, addresses three key challenges in generative information retrieval (GIR): vulnerability to prefix-level errors and local optima in constrained left-to-right decoding, the largely unimodal nature of prior GIR research, and the accuracy gap between discrete identifier-based GIR and dense-vector-based retrieval methods. DrIG aims to overcome these limitations by employing dual-role identifiers, which serve both as generation targets and as retrieval cues, enabling efficient and accurate retrieval across multiple modalities and domains. The framework is designed to handle instruction-aware retrieval, allowing users to specify the type of item they seek (text, image, or mixed). The paper's abstract highlights the potential of DrIG to advance GIR by improving accuracy while maintaining efficiency. The research is part of ongoing efforts to enhance information retrieval systems, particularly in the context of multimodal data. The paper is available at the provided arXiv URL.
Key facts
- Paper: arXiv:2608.12987
- Title: Generative Universal Multimodal Retrieval with Dual-role Identifiers
- Introduces DrIG framework
- Supports retrieval across text, image, and mixed image-text items
- Addresses three challenges: prefix-level errors, unimodal research, accuracy gap
- Uses dual-role identifiers for generation and retrieval
- Announce type: cross
- Published on arXiv
Entities
Institutions
- arXiv