Search papers, labs, and topics across Lattice.
This paper introduces $\omega$-0, a latent predictive world-action model designed for humanoid robots to perform concurrent loco-manipulation tasks by predicting whole-body action latents from language instructions and sensory inputs. By integrating visual foresight with a diffusion-based action generation approach, $\omega$-0 enables real-time, coordinated movements that are essential for complex household tasks. Experimental results show that $\omega$-0 significantly outperforms existing methods, achieving smoother and more effective manipulation while moving across various real-world scenarios.
$\omega$-0 enables humanoid robots to seamlessly integrate movement and manipulation, outperforming traditional models by predicting coordinated actions directly from sensory inputs.
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $\omega$-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, $\omega$-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, $\omega$-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect $\omega$-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single $\omega$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.