Search papers, labs, and topics across Lattice.
This paper introduces RoboReact, a novel framework that synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation by generating human manipulation videos and extracting keyframes through depth-aware 3D reconstruction. The framework effectively bridges the gap between imagined plans and physical execution by employing online object-centric re-grounding and a vision-language model-guided refinement loop, allowing for robust adaptation to geometric mismatches and execution deviations. Experiments on real humanoid robots show that RoboReact can generalize across various object configurations and recover from disturbances without the need for teleoperation or human demonstrations, underscoring its potential for scalable skill acquisition in robotics.
RoboReact enables humanoid robots to learn dexterous manipulation skills from synthetic video data, achieving robust performance without human intervention.
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.