Search papers, labs, and topics across Lattice.
This paper introduces Hand2Bot, a novel RGB-D video dataset designed to enhance human-to-robot object handover prediction by capturing rich contextual cues like body posture and facial expressions in realistic environments. The authors also propose PassGen, a generative pipeline that synthesizes handover sequences using stable video diffusion and an Intention-Aware Temporal Face Encoder, while addressing the sim-to-real gap through a morphology-based depth editing strategy. Experimental results show that training with PassGen significantly improves intention identification accuracy and reduces false triggers, enabling robots to better anticipate human actions in collaborative settings.
Training on the PassGen framework allows robots to anticipate human intentions earlier and more accurately, transforming human-robot collaboration.
Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and the significant sim-to-real gap. To address these challenges, we introduce Hand2Bot, an RGB-D video dataset that provides rich contextual information such as body posture and facial expressions, specifically collected for handover scenarios with real-world noise patterns. We further propose PassGen, a generative pipeline that leverages stable video diffusion and an Intention-Aware Temporal Face Encoder to synthesize realistic handover sequences while ensuring hand-object consistency. To bridge the sim-to-real gap, we implement a morphology-based depth editing strategy that replicates realistic sensor noise found in physical depth maps. Experimental evaluations demonstrate that our framework achieves high intention identification accuracy and low false trigger rates in both ablation studies and real-world deployment on a physical robot platform. Our results confirm that training on PassGen allows for robust zero-shot transfer and earlier intention anticipation compared to traditional hand-centric baselines, effectively enabling socially aware robotic behavior in shared workspaces.