Search papers, labs, and topics across Lattice.
The ADEPT framework leverages a two-phase reinforcement learning approach to enhance dexterity in high degree-of-freedom robots by first pretraining on a generic object reposing task and then post-training on specific downstream tasks. This method enables the robots to zero-shot the reposing phase of new tasks, significantly improving their ability to learn complex behaviors without starting from scratch. Key results demonstrate that ADEPT achieves human-level dexterity and speed in solving long-horizon tasks across different robotic embodiments, while maintaining stability during transfer learning through a novel post-training strategy.
Robots can now learn complex dexterous tasks at human-level speed without repetitive skill training, thanks to a novel pre-training and post-training reinforcement learning approach.
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies with this pretrained behavior as a prior. ADEPT enables learning new behaviors that are otherwise difficult to discover from scratch on multi-fingered robots and avoids learning the same set of skills over again for every new downstream task. The pretrained policy zero-shots the reposing phase of downstream tasks, but na\"ive RL fine-tuning rapidly degrades this capability during transfer. We address this with a stable post-training recipe combining behavior-cloning distillation, critic warm-up, and conservative on-policy updates. To safely exploit the full kinematic dexterity, we introduce a joint-space Geometric Fabric that mediates between the RL policy and the robot. We distill post-trained teachers into perceptive students that zero-shot sim-to-real transfer on two embodiments: a 23 DoF Kuka-Allegro with two RGB cameras, and a 29 DoF Flexiv-Sharpa with two RGB cameras and five vision-based tactile sensors, and can solve long-horizon tasks from challenging initial states with dexterity at human-level speed.