Search papers, labs, and topics across Lattice.
This paper introduces a unified framework that enhances real-world online reinforcement learning for robotic manipulation by integrating centralized training with decentralized execution and a Hybrid Reward Architecture. By decomposing the critic into task and grasp heads, the approach effectively addresses the challenges of non-stationarity and sample inefficiency associated with concurrent agent training. Experimental results show significant improvements in policy performance and sample efficiency, achieving success rates of up to 95% in complex tasks, far exceeding previous benchmarks.
Achieving up to 95% success in robotic manipulation tasks, this framework redefines the boundaries of sample efficiency and policy performance in real-world online reinforcement learning.
Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, avoiding the sim-to-real gap and enabling continuous policy refinement through human-in-the-loop interaction. Recent methods have demonstrated sample-efficient learning through human intervention but remain limited to small randomization ranges and encounter challenges with the non-stationarity induced by concurrently training multiple agents. To address these limitations, we introduce a unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA). This enables multiple actors to share a centralized multi-head critic. The critic is decomposed into task and grasp heads, corresponding to the sparse task reward and a potential-based grasping reward, respectively. We accordingly reformulate the critic and actor objectives to exploit the decomposed Q-values while explicitly accounting for the categorical action distribution of the discrete gripper policy. Experimental results demonstrate that the proposed framework substantially improves both sample efficiency and policy performance. We validate our approach on two robotic arms and a simulated humanoid robot across tennis ball and banana pick-and-place, pot reset, and simulated block relocation tasks under dimension-wise domain randomization, approximately 5-25x larger than those considered in prior work. Compared with a state-of-the-art baseline, our method improves the success rate from 60% to 80% on tennis ball pick-and-place, from 60% to 90% on banana pick-and-place, and from 25% to 95% on simulated block relocation, while also successfully accomplishing a task where the baseline consistently fails. Videos and more details are available at our project website: https://hil-harc.github.io/.