Search papers, labs, and topics across Lattice.
This paper introduces a model-free reinforcement learning framework tailored for continuous-time extended mean field control problems, addressing the challenges posed by joint state and control distributions. By employing deterministic feedback policies, the authors circumvent the complexities of optimizing stochastic kernels, leading to a novel deterministic policy gradient formula based on an advantage-rate function in Wasserstein space. The proposed method is validated through numerical experiments, showcasing its efficiency and robustness in scenarios with explicit control distribution dependencies.
Deterministic policies can significantly enhance the stability and efficiency of reinforcement learning in complex mean field control problems, outperforming traditional stochastic approaches.
This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls. We adopt deterministic feedback policies, under which the state--action distribution is induced directly as a push--forward of the state law. This avoids optimization over stochastic kernels and bypasses key limitations of existing approaches in extended mean field settings. We first establish a model--free sensitivity formula for parameterized McKean--Vlasov dynamics and use it to derive a deterministic policy gradient formula expressed through an advantage--rate function on the Wasserstein space. We then refine this formula by introducing local value and advantage--rate representations that depend on the state, action, and joint state--action distribution, yielding a policy gradient that includes both action derivatives and measure--derivative terms with respect to the control distribution. These characterizations lead to a martingale--based learning principle and motivate a continuous--time deep deterministic policy gradient algorithm combining particle approximations, measure--dependent neural networks, temporal--difference learning, and exploration in either action or parameter space. Numerical experiments on stochastic Cucker--Smale consensus control and optimal liquidation with trade crowding demonstrate the efficiency, stability, and robustness of the proposed method, including problems with explicit dependence on the control distribution.