H+ Embedding: A New Retrieval Model for Terminology-Intensive Domains
A new multi-granularity retrieval model called H+ Embedding has been developed by researchers to tackle the difficulties associated with terminology-heavy retrieval, especially in medical and scientific contexts. This model forecasts variable-length phrase segments, treats uncovered tokens as individual entities, and utilizes importance-guided unit selection through weighted MaxSim interaction. Its goal is to reconcile the limitations of single-vector retrievers, which tend to overly compress local relevance signals, with token-level late interaction techniques that lead to high indexing and storage expenses. In tests involving 16 scientific, medical, and bilingual tasks, H+ Embedding's phrase retrieval component surpassed the global retrieval component by 6.91 macro nDCG@10, closely approaching the performance of pricier token-level methods. The research paper can be accessed on arXiv with the identifier 2608.00065.
Key facts
- H+ Embedding is a unified multi-granularity retriever.
- It predicts variable-length phrase partitions and preserves uncovered tokens as singletons.
- It uses importance-guided unit selection with weighted MaxSim interaction.
- Evaluated on 16 scientific, medical, and bilingual tasks.
- Phrase retrieval branch exceeds global retrieval branch by 6.91 macro nDCG@10.
- Aims to balance retrieval accuracy and computational cost.
- Paper available on arXiv:2608.00065.
Entities
Institutions
- arXiv