Search papers, labs, and topics across Lattice.
This paper introduces a novel evolutionary policy optimization framework for scheduling heterogeneous agile Earth observation satellites, addressing the complexities of task selection and resource constraints across different platforms. By integrating reinforcement learning to guide operator selection within a memetic evolutionary algorithm, the approach enhances both global exploration and local refinement while maintaining feasibility under limited evaluations. Experimental results demonstrate that the proposed method, RLOSMEA, outperforms traditional metaheuristic baselines in terms of overall utility and convergence stability across various scheduling scenarios.
Reinforcement learning can significantly boost the efficiency of scheduling heterogeneous satellites, achieving better utility and convergence than traditional optimization methods.
Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage constraints. Since satellites differ in orbital access, maneuvering capability, and payload resources, the same task may have different feasible windows, transition costs, and resource-consumption patterns on different platforms, which increases the difficulty of unified modeling and efficient optimization. To address this problem, this paper proposes an evolutionary policy optimization framework for heterogeneous AEOS scheduling with preference-adjustable weighted objectives. In the modeling layer, assignment-based indirect encoding is combined with decoder-based equivalent-cost evaluation to retain satellite-dependent constraints while integrating task gain, energy saving, and load balance into an interpretable scalar utility. In the optimization layer, schedule decoding, population-based search, and online actor-critic operator control are decoupled, so that reinforcement learning selects high-level search operators rather than constructing schedules directly. Based on this framework, a reinforcement-learning-assisted operator-selection memetic evolutionary algorithm (RLOSMEA) is developed to coordinate global exploration, feasibility recovery, and local refinement under a limited function-evaluation budget. Experiments on different heterogeneous AEOS scenarios show that RLOSMEA achieves higher overall weighted utility and more stable convergence than representative metaheuristic baselines. Sensitivity and learning-behavior analyses further confirm the robustness of the proposed method and the effectiveness of reinforcement-learning-guided operator selection.