Search papers, labs, and topics across Lattice.
This paper introduces Proximal Prior Injection (ProxPI), a method that enhances model predictive control (MPC) by integrating learned policies while addressing the challenges posed by learned-prior mismatches. By employing a soft proximity cost, ProxPI maintains nominal-centered sampling, allowing the optimizer to escape unsuitable policy outputs and recover optimal performance akin to vanilla MPPI. Experimental results on both simulations and real robots confirm that ProxPI achieves robust performance across in-distribution and out-of-distribution tasks, outperforming traditional prior-injection methods.
ProxPI enables robust recovery from policy mismatches, ensuring that learned priors enhance rather than hinder performance in MPC scenarios.
Combining learned policies with model predictive control can leverage learned task priors while retaining online adaptation to new objectives and constraints, but performance degrades when the policy is out of distribution. In policy-guided model predictive path integral (MPPI) control, a policy-centered warm-start approach centers the sampling distribution on the policy output. When the prior is mismatched, centering the sampling distribution on the policy output restricts exploration around an unsuitable solution and prevents recovery toward the task optimum. We propose Proximal Prior Injection (ProxPI), which retains nominal-centered MPPI sampling and incorporates the policy through a soft proximity cost. This matches the in-distribution performance of existing prior-injection schemes while enabling the optimizer to escape an inaccurate policy and recover vanilla MPPI-level performance. We theoretically show that re-centering on the prior discards the optimizer's correction at every update, whereas nominal-centered sampling retains it and converges to a solution set by both the task cost and the prior, and that this failure is not removed by a larger rollout budget. Simulations and real-robot experiments demonstrate robust performance under both in-distribution and out-of-distribution tasks.