GOTS Method Reduces Visual Tokens in High-Resolution VLMs
A novel approach known as Greedy Orthogonal Token Selection (GOTS) eliminates the need for training while minimizing the number of visual tokens in high-resolution vision-language models (VLMs). Contemporary VLMs utilize dynamic or high-resolution visual encoding, producing thousands of tokens that elevate inference costs. Current reduction techniques evaluate token utility based on importance, relevance, coverage, diversity, or subset-level goals. In contrast, GOTS focuses on selected-span complementarity by choosing the token with the highest residual energy that is orthogonal to the currently retained subset, thereby optimizing the one-step augmented Gram determinant. This method is query-agnostic and does not require any training. The research can be found on arXiv with ID 2607.23913.
Key facts
- GOTS is a training-free and query-agnostic method for visual token reduction.
- It selects tokens based on orthogonal complementarity to the retained subset.
- The method maximizes the one-step augmented Gram determinant.
- Modern VLMs produce thousands of visual tokens, increasing inference cost.
- Existing methods use token-wise importance, query relevance, coverage, or diversity.
- The paper is published on arXiv with ID 2607.23913.
- GOTS does not require training or query information.
- The approach views token reduction through selected-span complementarity.
Entities
Institutions
- arXiv