Search papers, labs, and topics across Lattice.
This paper introduces the World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework to enhance collaborative task offloading in Low Earth Orbit (LEO) satellite networks, addressing the challenges posed by intermittent inter-satellite links and dynamic traffic loads. By formulating the offloading problem as a partially observable multi-agent decision-making process, the framework leverages a predictive world model to enable proactive planning and a Transformer-based policy architecture to capture inter-agent dependencies. The results demonstrate that WM-MAPPO significantly outperforms traditional model-free MARL methods and other heuristic approaches in terms of task completion ratios, latency, energy efficiency, and robustness.
WM-MAPPO achieves superior task completion and efficiency in LEO satellite networks by leveraging a predictive world model and advanced inter-agent coordination.
Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.