RIFT: Rollout-Free Imagination via Future Tokens for World Action Models
A recent study published on arXiv (2608.11521) examines whether world action models (WAMs) need the complete iterative video rollout trajectory for robot action generation or if a static future representation is adequate. Analyzing four WAMs across all 40 LIBERO tasks, the research indicates that while action generation is influenced by future-cache values and their locations, specific models (Joint and Cosmos-2) can effectively utilize a fixed final-clean key/value (K/V) cache with only slight performance degradation, achieving success rates between 97.9% and 98.2% and an end-effector average displacement error of 1.7 to 1.9 cm. This discovery differentiates cache consumption from production, leading the authors to introduce RIFT (Rollout-free Imagination via Future Tokens), which leverages learned anticipation to create future tokens without iterative rollout, potentially minimizing deployment latency. The paper is a cross-type submission by the research team on arXiv.
Key facts
- The paper is titled 'Keep the Future, Drop the Rollout: RIFT for World Action Models'.
- It is available on arXiv with ID 2608.11521.
- The study tested four WAMs on all 40 LIBERO tasks.
- Masking or reassigning future-cache values changes execution and reduces success.
- Joint and Cosmos-2 models can reuse a fixed final-clean K/V cache with 97.9% to 98.2% success.
- End-effector average displacement error is 1.7 to 1.9 cm.
- The proposed method is called RIFT (Rollout-free Imagination via Future Tokens).
- RIFT uses learned anticipation to avoid iterative rollout.
Entities
Institutions
- arXiv