Search papers, labs, and topics across Lattice.
The Kinematic-Aware Articulation Interface (KAI) introduces a structured representation that captures the kinematic structure of articulated objects, enhancing the efficiency of policy learning in robotic manipulation. By embedding geometric and kinematic priors, KAI significantly improves sample efficiency, achieving an average success rate of 82.9% across six simulation tasks while utilizing only half the demonstration data compared to traditional methods. Additionally, KAI demonstrates strong generalization capabilities, successfully transferring learned skills from controlled environments to complex real-world scenarios with visual distractions.
Achieving 82.9% success in articulated object manipulation with only half the data, KAI redefines efficiency in robotic learning.
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.