Search papers, labs, and topics across Lattice.
3
1
5
7
Current action-conditioned world models are limited by their reliance on visual patterns, failing to generalize physical dynamics across different robot embodiments.
Agents can now seamlessly transition between seeking and following tasks in dynamic environments, setting a new benchmark for embodied AI performance.
Ditch the multi-camera setup: VERM leverages foundation models to synthesize a single, task-optimized "virtual eye" view for robots, slashing training time by 1.89x.