ARTFEED — Contemporary Art Intelligence

GOTS Method Reduces Visual Tokens in High-Resolution VLMs

ai-technology · 2026-07-29

A novel approach known as Greedy Orthogonal Token Selection (GOTS) eliminates the need for training while minimizing the number of visual tokens in high-resolution vision-language models (VLMs). Contemporary VLMs utilize dynamic or high-resolution visual encoding, producing thousands of tokens that elevate inference costs. Current reduction techniques evaluate token utility based on importance, relevance, coverage, diversity, or subset-level goals. In contrast, GOTS focuses on selected-span complementarity by choosing the token with the highest residual energy that is orthogonal to the currently retained subset, thereby optimizing the one-step augmented Gram determinant. This method is query-agnostic and does not require any training. The research can be found on arXiv with ID 2607.23913.

Key facts

  • GOTS is a training-free and query-agnostic method for visual token reduction.
  • It selects tokens based on orthogonal complementarity to the retained subset.
  • The method maximizes the one-step augmented Gram determinant.
  • Modern VLMs produce thousands of visual tokens, increasing inference cost.
  • Existing methods use token-wise importance, query relevance, coverage, or diversity.
  • The paper is published on arXiv with ID 2607.23913.
  • GOTS does not require training or query information.
  • The approach views token reduction through selected-span complementarity.

Entities

Institutions

  • arXiv

Sources