Search papers, labs, and topics across Lattice.
This paper addresses the challenges of sample efficiency and numerical stability in sampling-based Model Predictive Control (MPC) for high-dimensional, open-loop unstable dynamical systems. By introducing Feedback Sampling MPC (FS-MPC), which optimizes the sampling proposal distribution through a hybrid design that balances local and global search strategies, the authors demonstrate significant improvements in convergence speed and optimality compared to standard methods. Empirical results on complex tasks, including humanoid loco-manipulation, confirm that FS-MPC outperforms traditional feedback policies and standard sample-based approaches in dynamically unstable environments.
FS-MPC achieves superior sample efficiency and stability in controlling high-dimensional robotic systems, outperforming traditional methods in challenging tasks.
Thanks to its parallelizability and flexibility, sampling-based Model Predictive Control (MPC) has become widely popular for controlling real-world robotic systems. However, for high-dimensional and open-loop unstable dynamical systems, the required number of samples to improve the control sequence will grow exponentially with the horizon, leading to poor sample efficiency and numerical instability. This paper investigates the instability of shooting methods in sampling-based MPC and shows that the optimal sampling proposal distribution can be realized by sampling with an optimized feedback policy. We refer to this algorithm as Feedback Sampling MPC (FS-MPC). FS-MPC involves a hybrid sampling design which balances local and global search based on the system stability and the available computation budget. Our theoretical analysis shows that our hybrid sampling approach achieves faster convergence than standard MPPI and better optimality than standard feedback sampling. Empirically, in diverse contact-rich control tasks like humanoid loco-manipulation and dexterous manipulation, we show that FS-MPC successfully tackles dynamically unstable tasks where standard sample-based approaches struggle, and strictly outperforms feedback policies alone. Finally, we validate our method on humanoid robot locomotion and manipulation tasks in the real world.