Search papers, labs, and topics across Lattice.
This paper introduces the Actuator Dynamics Curriculum, a novel approach that adjusts joint stiffness in legged robots to enhance learning in narrow-viability tasks where exploration is hindered. By initializing joint stiffness at a high value and gradually reducing it, the method increases the viability kernel of the Markov Decision Process, allowing for more effective exploration and training. The approach is validated through a cart-pole example and successfully applied to a quadrupedal-to-handstand transition on the Boston Dynamics Spot, demonstrating significant improvements in task execution both in simulation and hardware transfer.
Adjusting joint stiffness dynamically can transform the training landscape for legged robots, enabling successful execution of complex tasks that standard methods fail to solve.
Reinforcement learning has produced capable controllers across a broad range of legged-robot tasks, but a subset of these tasks fail to converge under standard training: those for which most exploration trajectories terminate before producing useful gradient signal. To address such tasks we introduce the \emph{Actuator Dynamics Curriculum}, a procedure that initializes joint stiffness at a high value and anneals it toward the system-identified value as completed episode lengths grow. Using a cart-pole system as a representative example, we show that higher closed-loop joint natural frequency under critical damping enlarges the viability kernel of the underlying Markov Decision Process, increasing the fraction of initial states from which the task is feasible. We validate the kernel monotonicity on the cart-pole and apply the curriculum to a quadrupedal-to-handstand transition on the Boston Dynamics Spot, a narrow-viability task where training under fixed identified stiffness plateaus at a policy that never completes the transition. The trained policy executes the transition in simulation across 10 seeds and transfers to hardware. More broadly, our results suggest that simulated actuator dynamics is a useful axis along which to design curricula for tasks in which exploration is bottlenecked by termination conditions rather than by reward signal.