Perturbation Grounded Selection: A New Approach to Vision-Language Test-Time Scaling
A new study available on arXiv (2608.01207) examines the reasons behind the underperformance of test-time scaling, a method that enhances reasoning in large language models by sampling and choosing from various candidate solutions, when utilized with vision-language models (VLMs). Earlier findings suggested that simple majority voting is more effective than selection methods reliant on the model’s self-verification, as answers grounded in images and confident guesses from language priors are often indistinguishable at the selection stage. To tackle this issue, the authors introduce Perturbation Grounded Selection (Pgs), a label-free and training-free approach that evaluates candidates based on their re-derivation under input perturbations like cropping or mild jitter. When no perturbations are applied, Pgs defaults to majority voting. The paper investigates whether Pgs surpasses chain-of-thought majority voting, likely providing empirical insights. This research advances the field of VLM reasoning enhancement.
Key facts
- arXiv paper 2608.01207
- Test-time scaling lifts LLM reasoning by sampling and selecting candidate solutions
- Simple majority voting beats self-verification selection methods for VLMs
- Pgs scores candidates by re-derivation under label-preserving perturbations
- Pgs recovers majority voting when perturbation set is empty
- Perturbations include cropping, background masking, photometric or geometric jitter
- Pgs is label-free and training-free
- The decisive question is whether Pgs beats chain-of-thought majority voting
Entities
Institutions
- arXiv