Search papers, labs, and topics across Lattice.
This paper introduces AERIS, an offline policy improvement framework designed for multi-UAV integrated sensing and communication (ISAC) that addresses the challenges of balancing communication quality, sensing reliability, and flight safety in dynamic environments. By utilizing fixed flight logs for centralized training and enabling decentralized execution, AERIS allows each UAV to make informed decisions based on local histories while leveraging global data for team-level assessments. The proposed STAR-CRDT algorithm significantly enhances the ISAC objectives, achieving a 29.3% improvement in return and notable gains in communication and sensing metrics while reducing collision risks by over half.
AERIS achieves a 29.3% boost in multi-UAV ISAC performance without the risks of trial-and-error learning.
Unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC) is a promising 6G paradigm, but dynamic multi-UAV ISAC control must jointly balance communication quality, sensing reliability, and flight safety under stochastic mobility. Existing optimization methods often require repeated global non-convex solving, while online reinforcement learning (RL) depends on risky trial-and-error flights that may cause sensing loss or collision-risk events. This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC. AERIS learns from fixed flight logs under centralized training and decentralized execution, so each UAV acts from local histories while training uses logged global information to assess team-level effects. We further design STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted improvements into the decentralized actor. We prove an offline-support policy improvement guarantee. Experiments show that STAR-CRDT improves the main ISAC objective return by 29.3% over the strongest baseline. It further improves communication sum rate, sensing pass rate, and sensing margin by 3.4%, 4.8%, and 69.1%, while reducing collision-risk events by 54.2%. On unseen real-road maps built from OpenStreetMap data, STAR-CRDT still obtains the best return.