Search papers, labs, and topics across Lattice.
PhiZero introduces a novel physical world model that leverages a compact discrete representation of world-state transitions through physical language, contrasting with traditional models that rely on pixel-space predictions. By learning this physical language from in-the-wild videos via self-supervision, PhiZero enables explicit reasoning about physical world dynamics, adopting a reason-then-render approach. Extensive experiments demonstrate its effectiveness in generating coherent world evolution and highlight its capabilities in realistic interactive modeling and zero-shot motion transfer.
PhiZero reveals that using physical language for world modeling can significantly enhance reasoning and simulation capabilities compared to traditional pixel-based methods.
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans'ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.