WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Efficient Autonomous Driving VLA
A new framework, WAM-Diff2, aims to combine the cognitive capabilities of autoregressive Vision-Language-Action (VLA) models with the execution efficiency of diffusion policies for autonomous driving. The approach addresses the challenge of transforming pre-trained autoregressive generalists into parallel diffusion models, which is difficult due to mismatched attention patterns and divergent optimization objectives. WAM-Diff2 is a multi-task discrete diffusion VLA framework that employs a three-stage hierarchical distillation process to bridge this gap. The paper, available on arXiv (2608.01035), details the architecture and methodology, highlighting the potential for low-latency, parallel execution while retaining holistic visual-linguistic reasoning. The work is significant for the deployment of VLA models in real-time autonomous driving systems, where computational latency and exposure bias are critical constraints.
Key facts
- WAM-Diff2 is a multi-task discrete diffusion VLA framework.
- It uses a three-stage hierarchical AR-to-diffusion distillation.
- The framework combines autoregressive VLA models with diffusion policies.
- It addresses computational latency and exposure bias in autonomous driving.
- The paper is available on arXiv with ID 2608.01035.
- The approach aims to transform pre-trained autoregressive generalists into parallel diffusion models.
- It tackles challenges of mismatched attention patterns and divergent optimization objectives.
- The framework is designed for end-to-end autonomous driving.
Entities
Institutions
- arXiv