Search papers, labs, and topics across Lattice.
This paper introduces WarpSAC, a family of off-policy reinforcement learning algorithms that adapts stabilizers based on the data regime, addressing the challenges posed by massively parallel simulation. Through extensive experiments across various benchmark families, the authors demonstrate that traditional stabilizers are ineffective in data-abundant environments, leading to the development of two variants: WarpSAC-L for data-limited scenarios and WarpSAC-A for data-rich contexts. WarpSAC achieves significant performance improvements, including a 4.5% increase in normalized score-step AUC over FlashSAC and a dramatic rise in success rates for specific tasks, highlighting the necessity of regime-aware approaches in scalable off-policy RL.
WarpSAC achieves up to a 96.4% success rate in complex tasks by dynamically adjusting its learning stabilizers based on data availability.
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.