Search papers, labs, and topics across Lattice.
This paper introduces a reinforcement learning framework for quadruped locomotion that incorporates biologically inspired gait parameters into the policy's action space. By outputting gait parameters and using gait-shaping rewards, the method encourages alignment between intended gait and execution, leading to more interpretable and physically plausible gaits. Experiments in simulation and on a large-scale hydraulic quadruped (BeTheX-Q) demonstrate faster convergence, robustness to disturbances, and interpretable gait adaptation across varying conditions compared to pure RL baselines.
Ditch black-box RL for quadruped locomotion: encoding gait parameters directly into the policy yields faster training, interpretable behaviors, and real-world robustness on a massive hydraulic robot.
Designing reward functions for reinforcement learning (RL)-based quadruped locomotion often requires extensive trial-and-error, limiting efficiency and interpretability. Lack of interpretability is particularly critical for large-scale hydraulic quadrupeds, where undetected unstable behaviors during deployment can cause significant mechanical damage. This letter presents a training framework that integrates biologically inspired gait parameters into RL policies, allowing robots to learn locomotion that is both system-aware and human-interpretable. The actor outputs gait parameters鈥攕uch as gait period, phase offset, stride length, foot clearance, duty factor, and base height鈥攚hich are coupled with gait-shaping rewards to encourage intent鈥揺xecution alignment and physically plausible gait patterns. The framework improves training efficiency and provides interpretable signals for monitoring policy behavior. In simulation, we show that locomotion follows the self-generated gait intents as soft constraints, and ablation results demonstrate faster convergence than a pure-RL baseline. We further present reward-term ablation and coefficient sensitivity analyses, indicating that performance is not driven by a single shaping term and is robust to moderate coefficient changes. We validate the approach on BeTheX-Q (over 1.8 m, 350 kg), demonstrating robust real-world locomotion across 2 and 4 km/h walking, stepping-stones, and external disturbances. Finally, gait-parameter analysis reveals interpretable trends that reflect locomotion intent and adaptation across varying conditions.