WorldClaw: Agentic Framework for Large-Scale 3D World Generation
A recently published research paper presents WorldClaw, an innovative framework designed for the generation of expansive, freely navigable 3D environments based on open-ended text prompts. This system tackles the complexities of ensuring global spatial coherence while providing rich local content and assets that can be edited and reused. Through planning agents, a text prompt is converted into a detailed specification encompassing regions, terrain, assets, materials, and spatial relationships. WorldClaw constructs a coherent terrain foundation using semantic layouts, generative materials, reusable assets, and a region-aware height field. For areas requiring intricate details, it creates terrain-conditioned compositions and reconstructs editable textured meshes. The findings, available on arXiv under identifier 2608.05248, highlight WorldClaw's capability to produce large-scale, editable 3D worlds, marking a significant advancement in scalable open-world generation for applications in gaming, simulation, and virtual environments.
Key facts
- WorldClaw is a fully agentic, coarse-to-fine framework for open-world 3D scene generation.
- It uses planning agents to translate text prompts into structured specifications of regions, terrain, assets, materials, and spatial relations.
- The framework builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field.
- For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement.
- Render-based agents refine terrain, objects, appearance, and contacts.
- The paper is available on arXiv with identifier 2608.05248.
- The announcement type is 'new'.
- The framework produces large-scale, editable 3D worlds across diverse open-world prompts.
Entities
Institutions
- arXiv