Search papers, labs, and topics across Lattice.
This paper introduces DCRL (Divide-and-Conquer RL), a novel approach to offline goal-conditioned reinforcement learning that addresses the challenges of long-horizon tasks by recursively decomposing trajectory segments into a balanced binary tree. By updating values from the leaves to the root, DCRL mitigates the issues of inaccurate short-range estimates and overestimation in value backups, leading to more accurate long-range value learning. Empirical results demonstrate that DCRL significantly outperforms existing flat offline GCRL methods, achieving a new state-of-the-art average score on challenging long-horizon tasks.
DCRL reduces worst-case bootstrap depth from linear to logarithmic, leading to a dramatic improvement in long-horizon offline goal-conditioned RL performance.
Scaling offline goal-conditioned reinforcement learning (GCRL) to long-horizon tasks is difficult because (1) long-range value learning depends on shorter-range estimates that may still be inaccurate, and (2) max-based value backups can amplify overestimation through repeated propagation. We propose DCRL (Divide-and-Conquer RL), which recursively decomposes each trajectory segment into a balanced binary tree and trains the values from leaves to root. Each parent is therefore updated only after its children, using an exact factorization of the observed route rather than selecting among noisy alternatives. Since this objective learns values along demonstrated routes that are not necessarily optimal, DCRL jointly propagates values across trajectories to discover shorter routes. Thanks to the balanced binary tree, DCRL reduces worst-case bootstrap depth from linear to logarithmic, and this shorter dependency structure empirically corresponds to much slower error accumulation. Across diverse goal-reaching tasks, DCRL substantially outperforms prior flat offline GCRL methods, and on the five most challenging long-horizon OGBench tasks, it improves the best prior average score from 55 to 64, surpassing all flat and hierarchical baselines.