ARTFEED — Contemporary Art Intelligence

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

ai-technology · 2026-08-17

A new framework called ForgeWM has been introduced to help build efficient video world models that require only a few action steps. It transforms a bidirectional action-conditioned video generator into streamlined models using methods like domain adaptation and teacher-forced causal training. The models operate effectively at denoising budgets of 1, 2, and 4 steps. Additionally, ForgeWM includes a dual-path deployment system that manages both latency-sensitive interactions and optional rollouts. This research addresses the challenge of aligning causal distillation with interactive world models, ensuring that keyboard states and mouse movements are in sync with compressed latent chunks during training and rollout. You can find the paper on arXiv with the identifier 2608.14022.

Key facts

  • ForgeWM is a progressive framework for few-step action-conditioned video world models.
  • It uses domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching.
  • The students operate at denoising budgets of 1, 2, and 4 steps.
  • It supports a dual-path deployment protocol.
  • The paper is on arXiv with ID 2608.14022.
  • It addresses challenges with discrete keyboard states and continuous mouse motion.
  • It extends causal distillation to interactive world models.
  • The framework transforms a bidirectional generator into few-step world models.

Entities

Institutions

  • arXiv

Sources