HTC-VLM: Hybrid Token Compression Boosts Vision-Language Model Efficiency
HTC-VLM, a new hybrid visual token compression framework, has been developed by researchers to mitigate the computational and memory demands associated with vision-language models (VLMs) that depend on numerous visual tokens. This framework, outlined in a paper on arXiv (ID: 2512.08240), distinguishes between semantics and appearance through two integrated pathways: a continuous pathway that retains detailed ViT patch features and a discrete pathway that utilizes MGVQ quantization for semantic anchors, represented by four tokens. These pathways merge into a 580-token hybrid sequence, which is then compressed into a single token via a disentanglement attention mask and a bottleneck. HTC-VLM maintains 87.2% average performance across seven benchmarks: GQA, VQAv2, MMBench, MME, POPE, SEED-Bench, and ScienceQA-Image, surpassing existing compression methods that struggle with the balance between continuous compression and discrete quantization. The paper was marked as a replace-cross type on arXiv, signifying a revised edition. This research enhances the efficiency of AI, particularly in multimodal models, by providing a well-rounded token compression strategy.
Key facts
- HTC-VLM is a hybrid visual token compression framework for vision-language models.
- It disentangles semantics and appearance through continuous and discrete pathways.
- The continuous pathway preserves ViT patch features; the discrete pathway uses MGVQ quantization with four tokens.
- The pathways are fused into a 580-token hybrid sequence and compressed into a single token.
- A disentanglement attention mask and a bottleneck are used for compression.
- Achieves 87.2% average performance retention across seven benchmarks.
- Benchmarks include GQA, VQAv2, MMBench, MME, POPE, SEED-Bench, and ScienceQA-Image.
- The paper is available on arXiv with ID 2512.08240.
Entities
Institutions
- arXiv