OpenVLA Failure Prediction: Internal Activations Signal Impending Errors Under Visual Shift
An arXiv preprint (2606.29699) explores the potential of internal activations in OpenVLA, a model that integrates vision, language, and action, to forecast failures resulting from shifts in visual distribution. The research maintains a fixed policy and captures one MLP activation for each step within the LIBERO-10 benchmark. Task success drops significantly from 57% to 17% due to occlusion. In the context of failed matched-reset trajectories, a logistic probe analyzing layer-16 activations yields an AUROC of 0.972 and an AUPRC of 0.352, whereas action disagreement only achieves an AUROC of 0.496. The occlusion-trained probe, without refitting, scores AUROC 0.689 on failed camera-jitter episodes. Despite strong retrospective discrimination, the layer-16 monitor averages 3.32 warning onsets per clean episode, revealing challenges for practical early warning systems. This study, authored by researchers, was published on arXiv on June 26, 2026.
Key facts
- OpenVLA is a vision-language-action policy.
- Visual shifts can cause OpenVLA to fail after initially plausible behavior.
- The study uses LIBERO-10 benchmark.
- Occlusion reduces task success from 57% to 17%.
- Layer-16 logistic probe achieves AUROC 0.972 and AUPRC 0.352 on failed matched-reset trajectories.
- Action disagreement achieves AUROC 0.496 on the same trajectories.
- Occlusion-trained probe reaches AUROC 0.689 on failed camera-jitter episodes without refitting.
- The layer-16 monitor averages 3.32 warning onsets per clean episode in calibration check.
Entities
Institutions
- arXiv