DURA: Diffusion-Based Unrestricted Robotic Attacks on VLA Models
A recent paper published on arXiv (2608.10393) presents DURA, an unrestricted robotic attack based on diffusion that targets Vision-Language-Action (VLA) models. These models are employed for robotic manipulation tasks, yet their resilience against adversarial attacks remains largely unexamined, posing risks of real-world damage. Current attacks depend on pixel-space alterations or require white-box access, leading to visible distortions and limitations in practical applications. DURA creates visually convincing adversarial patches suitable for both white-box and black-box scenarios, with the latter needing only the victim model's predicted actions. By optimizing the latent trajectory of a pretrained diffusion model, DURA produces natural patches that direct the robot towards actions specified by the attacker. The study includes extensive simulations and (presumably) real-world experiments, underscoring vulnerabilities in VLA models and the necessity for stronger defenses in robotic systems.
Key facts
- DURA is a diffusion-based unrestricted robotic attack for VLA models.
- It generates visually natural adversarial patches.
- Supports white-box and black-box attack settings.
- Black-box setting requires only predicted actions.
- Optimizes along latent trajectory of pretrained diffusion model.
- Steers robot toward attacker-specified target actions.
- Experiments conducted in simulation and possibly physical environments.
- Paper available on arXiv with ID 2608.10393.
Entities
Institutions
- arXiv