ARTFEED — Contemporary Art Intelligence

HTC-VLM: Hybrid Token Compression Boosts Vision-Language Model Efficiency

ai-technology · 2026-08-13

HTC-VLM, a new hybrid visual token compression framework, has been developed by researchers to mitigate the computational and memory demands associated with vision-language models (VLMs) that depend on numerous visual tokens. This framework, outlined in a paper on arXiv (ID: 2512.08240), distinguishes between semantics and appearance through two integrated pathways: a continuous pathway that retains detailed ViT patch features and a discrete pathway that utilizes MGVQ quantization for semantic anchors, represented by four tokens. These pathways merge into a 580-token hybrid sequence, which is then compressed into a single token via a disentanglement attention mask and a bottleneck. HTC-VLM maintains 87.2% average performance across seven benchmarks: GQA, VQAv2, MMBench, MME, POPE, SEED-Bench, and ScienceQA-Image, surpassing existing compression methods that struggle with the balance between continuous compression and discrete quantization. The paper was marked as a replace-cross type on arXiv, signifying a revised edition. This research enhances the efficiency of AI, particularly in multimodal models, by providing a well-rounded token compression strategy.

Key facts

  • HTC-VLM is a hybrid visual token compression framework for vision-language models.
  • It disentangles semantics and appearance through continuous and discrete pathways.
  • The continuous pathway preserves ViT patch features; the discrete pathway uses MGVQ quantization with four tokens.
  • The pathways are fused into a 580-token hybrid sequence and compressed into a single token.
  • A disentanglement attention mask and a bottleneck are used for compression.
  • Achieves 87.2% average performance retention across seven benchmarks.
  • Benchmarks include GQA, VQAv2, MMBench, MME, POPE, SEED-Bench, and ScienceQA-Image.
  • The paper is available on arXiv with ID 2512.08240.

Entities

Institutions

  • arXiv

Sources