Search papers, labs, and topics across Lattice.
This paper introduces RoboEdit, a video editing suite that converts human manipulation videos into robot-compatible formats, enabling scalable training for robotic systems. By leveraging an automatic pipeline called RoboEdit-ADC, the authors generate a large dataset of 174K aligned video pairs, which includes diverse scenes and interaction types across seven robot embodiments. The results demonstrate that RoboEdit not only achieves superior editing quality but also enhances the performance of robot control policies in real-world tasks, showcasing its potential to harness abundant human video data for robotic learning.
RoboEdit transforms abundant human manipulation videos into high-fidelity training data for robots, unlocking new avenues for scalable robot learning.
Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning.