Search papers, labs, and topics across Lattice.
This paper introduces WANDA, a synthetic data engine that enables the learning of open-world mobile manipulation policies from a single human demonstration. By reconstructing environments and robot-object interactions, WANDA generates diverse trajectories and enhances long-horizon robustness through Corrective State Expansion, ultimately allowing for effective cross-environment generalization. Experimental results demonstrate that policies trained with WANDA achieve significant improvements in spatial generalization and robustness, even enabling zero-shot deployment on different robotic embodiments.
Achieving long-horizon robustness and spatial generalization from just one human demonstration could revolutionize how we approach data efficiency in robotic learning.
Learning open-world mobile manipulation policies requires vast data to achieve spatial generalization, long-horizon robustness, and scene generalization. Current prevailing data collection paradigms, teleoperation and UMI, demand prohibitive human effort and cost at scale. To scale beyond the limits of manual data collection, we seek to maximize the value of each human demonstration by scalable data generation. To this end, we introduce WANDA: learning open-World mobile mANipulation from one demonstration via a synthetic DAta engine. WANDA first reconstructs background Gaussian splats and robot-object interaction trajectories from source RGBD observations, as a world substrate for later planning and rendering. It then rearranges contact-rich robot-object interaction segments into extensive spatial configurations, utilizing whole-body motion planning to chain them into new trajectories. To enhance long-horizon robustness, it applies Corrective State Expansion to increase the robot and object state diversity at different stages of mobile manipulation. To unlock cross-environment generalization, trajectories are synthesized on diverse generated 3D worlds from everyday photos. Furthermore, we synthesize photo-realistic observations by compositing rendered robot and object meshes with Gaussian splatting backgrounds. We evaluate our approach on extensive simulation and real-world tasks in various scenes. Experiments show that policies trained with WANDA achieve long-horizon robustness, broad spatial generalization and cross-environment generalization from one real demonstration. Moreover, WANDA naturally supports cross-embodiment data generation, validated by zero-shot deployment on another mobile manipulator with a distinct morphology.