CRUISE: VLM-Guided Uncertainty-Aware Sensor Fusion for Autonomous Driving
A research paper named 'CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving' has been released on arXiv (arXiv:2608.09202v1). It tackles the issue of effective cross-modal feature fusion in self-driving cars, which depend on various sensors like cameras, LiDAR, and radar for understanding their surroundings. The performance of these sensors can greatly differ in various real-world scenarios, such as low visibility and harsh weather. Although uncertainty quantification (UQ) can help models focus on dependable signals, current methods often depend on basic feature-level uncertainty assessments that struggle in complex out-of-distribution situations. To address this, the authors introduce CRUISE, a framework that utilizes a vision-language model (VLM)-guided UQ module for detailed, pixel-level uncertainty estimates, enhancing sensor fusion reliability in difficult environments. This paper is classified as a new submission on arXiv.
Key facts
- Paper title: CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
- Published on arXiv with ID 2608.09202v1
- Addresses cross-modal sensor fusion in autonomous vehicles
- Sensors include cameras, LiDAR, and radar
- Challenges include poor visibility and adverse weather
- Existing uncertainty-aware fusion methods use simple feature-level estimates
- CRUISE integrates a VLM-guided UQ module for pixel-level uncertainty
- Leverages VLM's rich prior knowledge
Entities
Institutions
- arXiv