ARTFEED — Contemporary Art Intelligence

HAM-RAG: A New Framework for Structure-Faithful Multimodal Generation

ai-technology · 2026-08-17

A recent study presents HAM-RAG, a Hierarchy-Aware Multimodal RAG framework aimed at enhancing structure-preserving interleaved generation. This framework tackles a significant drawback of current multimodal RAG approaches, which typically reduce structured documents to separate text and image components, compromising the integrity of source organization and the necessary local text-image coherence for accurate evidence selection and placement. By utilizing document hierarchy as a grounding signal for both retrieval and generation, HAM-RAG maintains the source position and local text-image relationships in the prompt while contextualizing textual and visual evidence. Additionally, the authors introduce HAM-Bench, a benchmark comprising datasets from Wukong, Wiki, arXiv, and Recipe, covering game walkthroughs, web pages, scientific papers, and recipe documents. HAM-RAG enhances the main multimodal average by 17.3% across various backbones compared to the strongest non-hierarchical baseline. Notably, on the Wukong dataset, it shows a 24.2% improvement in Img-CBS over the strongest non-hierarchical baseline, indicating significantly better performance. The paper can be found on arXiv with the identifier 2608.14032.

Key facts

  • HAM-RAG is a Hierarchy-Aware Multimodal RAG framework.
  • It addresses limitations of existing multimodal RAG methods that flatten structured documents.
  • HAM-RAG uses document hierarchy as a grounding signal for retrieval and generation.
  • HAM-Bench covers Wukong, Wiki, arXiv, and Recipe datasets.
  • HAM-RAG improves the main multimodal average by 17.3% over the strongest non-hierarchical baseline.
  • On Wukong, HAM-RAG improves Img-CBS by 24.2% over the strongest non-hierarchical baseline.
  • The paper is available on arXiv with identifier 2608.14032.

Entities

Institutions

  • arXiv

Sources