Search papers, labs, and topics across Lattice.
This paper introduces Prefix-Optimal Generative Policies (POGP), a novel framework that optimizes the denoising process in diffusion policies for continuous control by learning a prefix value function at each denoising step. By employing a Bellman-style recursion, POGP significantly reduces the number of necessary denoising iterations by approximately 2.7-fold while maintaining near-full task performance across four MuJoCo environments. Additionally, the framework enhances final task performance by about 3.5% compared to state-of-the-art dynamic diffusion methods, demonstrating the dual benefits of supervised intermediate outputs for both early stopping and policy improvement.
Reducing denoising iterations by 2.7 times while boosting performance by 3.5% reveals a new frontier in optimizing diffusion policies for continuous control.
Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requires adapting the number of denoising steps to the difficulty of each action while preserving task performance. We introduce Prefix-Optimal Generative Policies (POGP), a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain. The prefix value function serves two purposes: it provides an auxiliary training objective that encourages intermediate outputs to become high-quality actions, and it enables a test-time stopping rule that terminates denoising when additional steps are unlikely to produce meaningful improvement. Across four MuJoCo environments and comparisons with 12 baselines, POGP reduces the required number of denoising iterations by approximately 2.7-fold while retaining near-full task performance. Compared with state-of-the-art dynamic diffusion baselines, prefix training also improves final task performance by approximately 3.5%. These results indicate that supervising intermediate denoising steps is useful not only for adaptive early stopping, but also as an auxiliary objective that improves the learned policy.