Search papers, labs, and topics across Lattice.
3
0
5
Video generative priors only translate into robust physical control when paired with explicit world-to-action information routing and synchronized joint denoising rather than standard monolithic fine-tuning.
Motion-aligned latent dynamics enable robots to learn actionable behaviors from human videos without losing visual fidelity across different embodiments.
Today's visually impressive video world models still fail to produce physically plausible robot actions, revealing a critical gap between visual realism and embodiment.