GUARD: Test-Time Failure Detection for Diffusion-Based Vision-Language-Action Policies
A novel method called GUARD has been developed to enhance the testing of diffusion-based vision-language-action (VLA) policies, which are capable of producing plausible actions despite weak connections to necessary visual and linguistic cues. GUARD meticulously analyzes the grounding of these actions while maintaining the original pretrained policy. It evaluates individual components in the model's key-value (KV) cache, generates alternative caches, and inspects variations in denoising outputs. Additionally, GUARD offers diagnostic insights such as sensitivity and attention entropy. This innovative technique was tested across five policy-benchmark scenarios, with findings available on arXiv under ID 2608.04510.
Key facts
- GUARD is a test-time failure detection method for diffusion-based VLA policies.
- It measures grounding without modifying the pretrained policy.
- It uses the final vision-language model key-value (KV) cache.
- It constructs counterfactual caches by ablating salient KV entries.
- It compares denoising responses with the original conditioning.
- It derives a diagnostic stream including sensitivity, attention entropy, modality bias, and grounding efficiency.
- The diagnostic stream is calibrated online and processed by a lightweight temporal classifier.
- Evaluation used Pi0, SmolVLA, and Alpamayo-1.5 on LIBERO, SimplerEnv, MetaWorld, and Physic.
Entities
Institutions
- arXiv