Search papers, labs, and topics across Lattice.
This paper introduces a novel safe offline multi-agent reinforcement learning algorithm that integrates neural individual control barrier functions into a diffusion model to enhance safety during trajectory generation. By addressing the safety challenges inherent in multi-agent environments, the method allows for the recovery of control policies through inverse dynamics, facilitating effective learning from offline data. The evaluation across various benchmarks shows significant safety improvements while still achieving competitive reward outcomes, highlighting the algorithm's efficacy in safety-critical applications.
Embedding individual control barrier functions into diffusion models can dramatically enhance safety in multi-agent reinforcement learning without sacrificing performance.
Offline reinforcement learning allows control policies to be learned directly from data without online interaction, making it suitable for safety-critical tasks. Recent studies have applied diffusion models to offline reinforcement learning to leverage their strong capacity for modeling complex data distributions. However, existing approaches primarily focus on single-agent settings, leaving the safety challenges in multi-agent environments largely unexplored. In this work, we propose a safe offline multi-agent reinforcement learning algorithm that embeds neural individual control barrier functions into the diffusion model to enhance safety during trajectory generation, with control policies recovered through inverse dynamics. We evaluate our algorithm across diverse benchmarks, demonstrating substantial safety improvements while maintaining competitive rewards.