Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
An arXiv paper (2603.06054) recently explores the shortcomings of Vision-Language Models (VLMs) when addressing straightforward visual inquiries pertinent to automated driving. The authors analyze the intermediate activations within VLMs to understand how particular visual concepts are represented linearly, seeking to pinpoint obstacles in the flow of visual information. They developed counterfactual image sets that vary solely in a specific visual concept and employed linear probes to differentiate them based on activations from five leading VLMs, including two variations of one model. Findings reveal that essential concepts, such as the existence of an object or agent, are inadequately encoded, resulting in failures. This research underscores the necessity for enhanced representation of visual concepts in VLMs for self-driving technology.
Key facts
- Paper arXiv:2603.06054, announced as replace-cross
- Focus on Vision-Language Models (VLMs) in automated driving
- Examines intermediate activations of VLMs
- Assesses linear encoding of specific visual concepts
- Uses counterfactual image sets differing only in targeted visual concept
- Trains linear probes on activations of five SOTA VLMs
- Includes two training variants for one model
- Results show concepts like presence of object or agent are not robustly encoded
Entities
Institutions
- arXiv