Search papers, labs, and topics across Lattice.
Hydra-0 is a novel generalist world model that leverages action flow by representing robot actions as pixel motion, facilitating a unified approach to world modeling and control. The model significantly outperforms its action-conditioned baseline, achieving a 90.4% reduction in robot-motion error and a 60.2% reduction in object-motion error, while also enabling zero-shot composition and efficient adaptation across various tasks and environments. Notably, it reveals an emergent capability to predict robot motion from human demonstrations, enhancing its utility in diverse robotic applications.
Action flow transforms robot control by allowing models to learn from diverse data without task-specific demonstrations, achieving unprecedented accuracy in motion prediction.
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.