Search papers, labs, and topics across Lattice.
ZimaBlue addresses the challenge of robust generalization in robotic manipulation by leveraging egocentric videos as a scalable source of embodied experience. The framework employs a three-stage training curriculum that includes causal embodied video pre-training, video-action mid-training, and specialization for target robots, ultimately enhancing the model's ability to predict actions in real-time. The results demonstrate a significant improvement in zero-shot evaluation success rates, from 36.1% to 77.8%, showcasing the effectiveness of large-scale video data in training generalizable World Action Models.
Scaling from 120,000 hours of egocentric video boosts robotic manipulation success rates from 36.1% to 77.8% in zero-shot evaluations.
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of embodied experience, capturing object interactions, contact dynamics, tool use, and long-horizon behaviors across diverse environments. The central challenge is how to convert this abundant but action-free experience into effective robot control. We introduce ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video. ZimaBlue follows a three-stage training curriculum: it first performs causal embodied video pre-training on large-scale human and robot egocentric videos, then grounds the learned visual dynamics in heterogeneous robot trajectories through video-action mid-training with a unified action representation, and finally specializes the model to a target robot for deployment. To make generative WAMs practical for real-time control, ZimaBluefurther adopts an asynchronous Slow-Fast dual-system architecture, where a high-capacity Slow world model provides generalizable spatiotemporal representations and a lightweight Fast branch enables 30 Hz action prediction on NVIDIA RTX 4090. On real-robot zero-shot evaluations, scaling from target-robot data alone to over 120,000 hours of embodied video improves success from 36.1% to 77.8%. ZimaBlue further delivers strong performance across multiple benchmarks, with particularly pronounced gains on unseen tasks.