Search papers, labs, and topics across Lattice.
WorldClaw introduces a novel coarse-to-fine framework for generating expansive 3D open worlds from text prompts, addressing the complexities of maintaining global spatial coherence and rich local detail. By employing planning agents to create structured specifications and utilizing a combination of generative and procedural techniques, the system constructs coherent terrains and detailed scenes that are also suitable for downstream editing. The key result demonstrates that WorldClaw can produce large-scale, visually compelling environments with editable assets, achieving a balance between spatial organization and local content richness across diverse prompts.
WorldClaw can generate expansive, editable 3D worlds from text prompts while maintaining both global coherence and intricate local details.
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.