Dual-purpose Semantic IDs boost LLM-level I/O efficiency in recommendation systems
A new research paper from arXiv proposes Dual-purpose Semantic IDs to overcome memory bottlenecks in large-scale recommendation systems. The method uses hierarchical quantization to convert dense embeddings into discrete tokens that serve both as collaborative identity for user-item interactions and as content reconstruction via a lightweight Semantic Decoder. This replaces massive vector storage with on-demand embedding approximation, reducing system overhead and data footprints. The framework was validated through offline evaluations and online deployment in production-scale ranking.
Key facts
- arXiv:2607.24865v1
- Announce Type: cross
- Proposes Dual-purpose Semantic IDs
- Uses hierarchical quantization
- Semantic IDs serve two roles: Collaborative Identity and Content Reconstruction
- Replaces massive vector storage with on-demand reconstruction
- Validated through offline evaluations and online deployment
- Aims to achieve LLM-level I/O efficiency
Entities
Institutions
- arXiv