Search papers, labs, and topics across Lattice.
This paper introduces Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation (EDPFRL-IM), which enhances Personalized Federated Reinforcement Learning (PFRL) by incorporating curiosity-driven exploration to improve policy discovery in sparse-reward environments. By utilizing intrinsic random network distillation (RND) signals alongside extrinsic rewards, the framework promotes local exploration while maintaining client data privacy. Experimental results demonstrate that EDPFRL-IM significantly outperforms existing PFRL benchmarks in terms of policy personalization and sample efficiency, particularly in challenging environments with delayed and sparse rewards.
Curiosity-driven exploration can dramatically enhance policy personalization in federated learning, even in sparse-reward scenarios.
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for each client, thereby neglecting exploration in non-stationary or sparse-reward environments. In this work, we introduce a new exploration-driven framework, Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation (EDPFRL-IM), that leverages an inherent curiosity-driven exploration at each client to promote local exploration and protect client privacy. Furthermore, to facilitate policy discovery via exploration in previously unexplored state spaces, clients add an intrinsic random network distillation (RND) signal to their extrinsic reward. Additionally, the server does not have access to clients'raw experiences or local gradient estimates; instead, the server sends global exploration priors and collects minimal novelty summaries from each client to enable both diverse and coordinated exploration among clients. Experiments in benchmark environments show that our framework outperforms average PFRL benchmarks in policy personalization and sample efficiency, primarily in delayed and sparse reward systems. Overall, EDPFRL-IM enables the integration of a flexible exploratory learning structure into federated reinforcement learning systems while preserving client privacy.