Shanghai InnovationSJTUApr 12, 2026arXiv:2604.10677

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment

Yifu Xu, Bokai Lin, Xinyu Zhan, Hongjie Fang, Yong-Lu Li, Lixin Yang

AI Summary

LIDEA, a novel imitation learning framework, addresses the challenge of transferring knowledge from human video demonstrations to robots by mitigating the embodiment gap. It employs a dual-stage transitive distillation pipeline to align human and robot representations in a shared latent space, coupled with an embodiment-agnostic alignment strategy for consistent 3D perception. Experiments demonstrate that LIDEA can replace up to 80% of robot demonstrations with human data and achieve improved out-of-distribution generalization by transferring unseen human patterns.

Key Contribution

Human videos can substitute for up to 80% of costly robot demonstrations in imitation learning, thanks to a new method for bridging the embodiment gap.

Abstract

Scaling up robot learning is hindered by the scarcity of robotic demonstrations, whereas human videos offer a vast, untapped source of interaction data. However, bridging the embodiment gap between human hands and robot arms remains a critical challenge. Existing cross-embodiment transfer strategies typically rely on visual editing, but they often introduce visual artifacts due to intrinsic discrepancies in visual appearance and 3D geometry. To address these limitations, we introduce LIDEA (Implicit Feature Distillation and Explicit Geometric Alignment), an imitation learning framework in which policy learning benefits from human demonstrations. In the 2D visual domain, LIDEA employs a dual-stage transitive distillation pipeline that aligns human and robot representations in a shared latent space. In the 3D geometric domain, we propose an embodiment-agnostic alignment strategy that explicitly decouples embodiment from interaction geometry, ensuring consistent 3D-aware perception. Extensive experiments empirically validate LIDEA from two perspectives: data efficiency and OOD robustness. Results show that human data substitutes up to 80% of costly robot demonstrations, and the framework successfully transfers unseen patterns from human videos for out-of-distribution generalization.

Computer Vision Multimodal Models Robotics & Embodied AI

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment

Related Papers