ARTFEED — Contemporary Art Intelligence

CoverPrune: Optimal Transport-Based Token Pruning for 3D VLMs

ai-technology · 2026-08-15

A recent paper published on arXiv (2608.13226) presents CoverPrune, a novel framework for token pruning in 3D Vision-Language Models (3D VLMs) that does not require training. This approach tackles the issue of high visual token counts that hinder performance during inference. Traditional token pruning techniques often prioritize diversity, leading to the elimination of key prototype tokens in favor of outliers, which disrupts the multi-view consistencies and geometric structures vital for effective spatial reasoning. CoverPrune shifts the focus from diversity maximization to maintaining visual evidence coverage, framing token pruning at inference as an Optimal Transport (OT) problem and introducing a Feature-Spatial-Temporal (FST) mechanism to address the complex subset selection challenge. The research, significant for enhancing the efficiency of 3D VLMs in spatial reasoning tasks, is credited to a team of researchers.

Key facts

  • CoverPrune is a training-free framework for token pruning in 3D VLMs.
  • It formulates token pruning as an Optimal Transport (OT) problem.
  • Existing diversity-based pruning methods discard representative prototype tokens.
  • CoverPrune aims to preserve visual evidence coverage.
  • The method uses a Feature-Spatial-Temporal (FST) mechanism.
  • The paper is available on arXiv with ID 2608.13226.
  • The announcement type is cross.
  • The work addresses computational bottlenecks in 3D VLMs.

Entities

Institutions

  • arXiv

Sources