Study Finds Vision-Language Models Fail Visual Grounding in Zero-Shot Control
A new preprint on arXiv, labeled 2608.06154, dives into how vision-language models (VLMs) make decisions when they're used as zero-shot controllers. The study involved various assessments, such as blind-image controls and checking for consistency with repeated inputs. They analyzed a whopping 32,874 decisions from nine direct-action models and six structured local VLMs across different setups and simulators. Findings showed that a constant-SLOW policy outperformed a scripted geometric controller, with many models either being image-invariant or showing minimal changes. Notably, models that detected hazards didn't adjust their LEFT and RIGHT responses when reflected. This underscores the necessity for robust visual grounding checks before using VLMs as controllers.
Key facts
- arXiv preprint 2608.06154v1 investigates visual grounding in zero-shot vision-language control.
- The study used an input-ablation battery including blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks.
- Analyzed 32,874 scored calls across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy.
- Experiments covered two embodiments and three simulators.
- A constant-SLOW policy outperformed a scripted geometric controller.
- Several models were image-invariant or nearly constant.
- Models that recognized longitudinal hazards still failed to transform LEFT and RIGHT under reflection.
- No local VLM met the joint longitudinal and lateral criteria.
Entities
Institutions
- arXiv