DLAM: Probabilistic Latent Action Model for Robot Learning
Researchers introduce DLAM (Distributional Latent Actions with Temporal Constraints), a novel approach for learning latent action models from action-free video to improve vision-language-action (VLA) models. The method addresses the scarcity of action-labeled robot data by leveraging abundant observational videos. DLAM models each transition as a diagonal Gaussian, grounding the mean in observed visual change via reconstruction conditioned on a reference frame. It applies normalized composition and reversal over equal-gap triplets to constrain both mean and variance, using a lightweight shared-correlation coefficient for variance composition. This probabilistic formulation avoids deterministic transition points that cause error propagation in recursive composition, a limitation of existing structured methods. The work is published as arXiv:2607.27138.
Key facts
- DLAM stands for Distributional Latent Actions with Temporal Constraints.
- It targets vision-language-action (VLA) models constrained by scarce action-labeled robot data.
- Action-free videos provide abundant observations of physical change.
- Latent action models extract priors from action-free data.
- Reconstruction-trained codes may lack structure for joint generation with robot actions.
- Existing structured methods add temporal constraints but retain deterministic transition points.
- Residual errors in locally inferred transitions can propagate under recursive composition.
- DLAM represents each transition as a diagonal Gaussian.
- Reconstruction conditioned on the reference frame grounds the mean in observed visual change.
- Normalized composition and reversal over equal-gap triplets constrain mean and dimension-wise variance.
- Variance composition uses a lightweight shared-correlation coefficient.
- The paper is available on arXiv with ID 2607.27138.
Entities
Institutions
- arXiv