Search papers, labs, and topics across Lattice.
This paper introduces P^3 (Probabilistic Policy Propagation), a novel optimization framework that addresses the variance and bias issues in VAE-based policy learning when using Proximal Policy Optimization (PPO). By coupling moment-based probabilistic methods with sampling-based calibration, P^3 significantly enhances data efficiency and reduces convergence steps in robotic learning tasks. Experimental results demonstrate that P^3 improves data efficiency from 64.6% to over 96% and accelerates convergence by more than 20%, showcasing its effectiveness in complex humanoid parkour environments.
P^3 transforms VAE-based policy learning by slashing data inefficiency and convergence time, achieving over 96% efficiency in challenging tasks.
Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy marginalizes over the latent distribution, whereas former implementations estimate its probability ratio and KL divergence using only one latent sample. We identify a fundamental but overlooked theoretical cause: naive single-sample approximations in stochastic latent space induce significant variance and bias in the surrogate loss. To address this, we introduce P^3 (Probabilistic Policy Propagation), a distribution-aware optimization framework for VAE-based policies. $P^3$ couples moment-based probabilistic method for stable and efficient learning with sampling-based calibration for robust policy behavior under latent uncertainty. In our experiments, P^3 boosts data efficiency from 64.6% to>96%, reduces convergence steps by>20%. Furthermore, P^3 is evaluated on challenging humanoid parkour tasks and shows an effective foundation for VAE-based PPO. Code is available at https://github.com/ylyem9x/P3_Open.