Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
1
Sampling-Guided Policy Search (SGPS), which couples recurring action-target refinement by sampling-based model-predictive control with first-order policy optimization with first-order policy optimization, and shows that refinement improves policy learning beyond initialization and tracking alone.
Visual RL training can be sped up by orders of magnitude using trajectory perturbations, achieving state-of-the-art results on MuJoCo with just a single RTX 4080.