ARTFEED — Contemporary Art Intelligence

H+ Embedding: A New Retrieval Model for Terminology-Intensive Domains

ai-technology · 2026-08-04

A new multi-granularity retrieval model called H+ Embedding has been developed by researchers to tackle the difficulties associated with terminology-heavy retrieval, especially in medical and scientific contexts. This model forecasts variable-length phrase segments, treats uncovered tokens as individual entities, and utilizes importance-guided unit selection through weighted MaxSim interaction. Its goal is to reconcile the limitations of single-vector retrievers, which tend to overly compress local relevance signals, with token-level late interaction techniques that lead to high indexing and storage expenses. In tests involving 16 scientific, medical, and bilingual tasks, H+ Embedding's phrase retrieval component surpassed the global retrieval component by 6.91 macro nDCG@10, closely approaching the performance of pricier token-level methods. The research paper can be accessed on arXiv with the identifier 2608.00065.

Key facts

  • H+ Embedding is a unified multi-granularity retriever.
  • It predicts variable-length phrase partitions and preserves uncovered tokens as singletons.
  • It uses importance-guided unit selection with weighted MaxSim interaction.
  • Evaluated on 16 scientific, medical, and bilingual tasks.
  • Phrase retrieval branch exceeds global retrieval branch by 6.91 macro nDCG@10.
  • Aims to balance retrieval accuracy and computational cost.
  • Paper available on arXiv:2608.00065.

Entities

Institutions

  • arXiv

Sources