Search papers, labs, and topics across Lattice.
Puffin-World unifies physical grounding, spatial simulation, and 3D generation into a single multimodal model without relying on external offline geometry pipelines. By explicitly modeling physics (gravity and latitude), geometry (depth), and appearance conditioned on a novel Omni-Camera representation, the framework propagates real-world physical dynamics across future frames. Trained on the new 16-million-sample Puffin-16M dataset, the architecture couples view synthesis with dense reconstruction to enable stable, closed-loop world exploration and simulation.
World models no longer require fragile offline reconstruction pipelines鈥攂aking native physics, depth, and camera pose directly into a unified multimodal generative process unlocks self-calibrating, closed-loop 3D spatial simulation at scale.
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.