Search papers, labs, and topics across Lattice.
This paper introduces RLMM-Flow, a novel flow-based framework for mobile manipulation that integrates expert flow-policy pretraining with latent-space reinforcement learning to enhance whole-body motion generation. By first learning a multimodal motion prior from expert demonstrations and then refining it through a latent steering network, the framework effectively balances exploration and exploitation in high-dimensional action spaces. Experimental results demonstrate that RLMM-Flow significantly outperforms traditional imitation-only methods and existing reinforcement learning approaches in terms of task success, collision avoidance, and trajectory quality, while maintaining efficient inference speeds.
RLMM-Flow achieves superior mobile manipulation performance by combining expert knowledge with adaptive latent-space optimization, surpassing traditional methods in both safety and efficiency.
Mobile manipulation requires generating whole-body action chunks that jointly satisfy goal reaching, collision avoidance, base kinematic constraints, manipulator joint limits, and trajectory smoothness. Flow-based generative policies provide an efficient paradigm for learning multimodal and temporally consistent motion priors from expert demonstrations, but imitation-only training cannot improve policy quality beyond the demonstration distribution. We propose RLMM-Flow, a flow-based mobile manipulation framework that combines expert flow-policy pretraining with latent-space reinforcement learning post-training. The framework first learns a flow policy that captures a multimodal whole-body motion prior from expert demonstrations. The pretrained flow policy is then frozen, while a latent steering network steers its initial noise toward higher-value action chunks. To stabilize high-dimensional latent optimization, we warm up an action-space critic before jointly training the latent critic and latent actor, and introduce coarse-to-fine latent steering that progressively expands control from a horizon-shared latent representation to a full-dimensional residual representation. Experiments on mobile manipulation motion-planning benchmarks show that RLMM-Flow substantially improves task success, collision avoidance, and trajectory quality over imitation-only flow policies and existing reinforcement learning post-training baselines, while preserving fast flow-based inference.