Marionette: A New World Model for Interactive Games with Explicit 3D State Prediction
Researchers have introduced Marionette, a novel world model for interactive games that explicitly predicts a 276-dimensional 3D world state, including multi-entity articulated skeletons, metric root trajectories, and rotations. Unlike traditional autoregressive models that generate visual observations directly, Marionette delegates exact geometric computation to a fixed, zero-parameter renderer, leaving the neural model to synthesize appearance. This approach aims to improve long-horizon consistency and controllability in game environments. The model is detailed in a paper on arXiv (2608.14530), announced as a cross-type submission. The work addresses the fragility of latent world property accumulation over time by separating state prediction from rendering, potentially offering a more robust framework for interactive simulations.
Key facts
- Marionette is a world model for interactive games with articulated characters.
- It predicts an explicit 276-dimensional 3D world state.
- The world state includes multi-entity articulated skeletons, metric root trajectories, and rotations.
- A zero-parameter graphics bridge converts predicted states into pose-control videos.
- The model uses a two-stage autoregressive dynamics model.
- It delegates geometric computation to a fixed renderer, leaving appearance synthesis to the neural model.
- The paper is available on arXiv with ID 2608.14530.
- The announcement type is cross.
Entities
Institutions
- arXiv