Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
A novel framework named Chain-of-Visual-Thought (COVT) has been unveiled to improve the perceptual capabilities of Vision-Language Models (VLMs). While VLMs are proficient in linguistic reasoning, they face challenges with tasks that demand intricate visual perception, like spatial reasoning and geometric understanding. This shortcoming stems from their inadequate ability to capture dense visual data across various spatial dimensions. COVT allows VLMs to reason with both language and continuous visual tokens—concise latent representations that convey rich perceptual information. With a modest allocation of around 20 tokens, COVT extracts insights from lightweight vision specialists, encompassing aspects like 2D appearance, 3D geometry, and edge structure. The framework is elaborated in a paper on arXiv (arXiv:2511.19418v3), which likely presents experimental findings on COVT's effectiveness, although specific performance metrics are not detailed in the abstract. This advancement is crucial for artificial intelligence, especially in computer vision and multimodal learning, as it addresses a recognized limitation in VLM functionality.
Key facts
- Chain-of-Visual-Thought (COVT) is a new framework for Vision-Language Models (VLMs).
- COVT enables VLMs to reason through continuous visual tokens, not just words.
- The framework uses roughly 20 tokens to distill knowledge from lightweight vision experts.
- COVT captures complementary properties: 2D appearance, 3D geometry, spatial layout, and edge structure.
- During training, the VLM autoregressively predicts visual tokens to reconstruct dense supervision signals.
- Supervision signals include depth, segmentation, edges, and DINO features.
- The paper is available on arXiv with identifier arXiv:2511.19418v3.
- The announcement type is 'replace-cross'.
Entities
Institutions
- arXiv