Search papers, labs, and topics across Lattice.
This paper introduces an implicit Q-learning-bootstrapped ant colony optimization method (IQACO) to tackle the complex problem of scheduling maritime moving-target observations with agile satellites. By integrating an offline implicit Q-learning module, IQACO dynamically adjusts key parameters of the ant colony optimization process, enhancing the scheduler's ability to adapt to varying observation windows and constraints. Experimental results across 14 scenarios reveal that IQACO consistently outperforms traditional ant colony optimization, achieving a 3.40% to 9.40% improvement in observation benefits while ensuring stability across different settings.
IQACO not only boosts observation efficiency for agile satellites but also adapts dynamically to the unpredictable nature of maritime targets.
Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, onboard-resource, and cloud-affected availability constraints. This paper proposes an implicit Q-learning-bootstrapped ant colony optimization method, termed IQACO, for multi-satellite maritime moving-target observation scheduling. Rather than directly learning a task-selection policy, IQACO embeds an offline implicit Q-learning module into constructive ant colony optimization to adaptively adjust the pheromone factor, heuristic factor, and evaporation rate. A compact search-state representation captures pheromone distribution, current and historical-best solution quality, and iteration progress. During online scheduling, ant colony optimization constructs feasible observation sequences, while the learned policy regulates exploration and exploitation according to the current search state. Experiments on 14 scenarios with different scales and satellite configurations show that IQACO obtains the highest mean observation benefit in every scenario, improves the result of conventional ant colony optimization by 3.40\%--9.40\%, accelerates convergence, and remains stable under different objective-weight settings. These results demonstrate that offline value learning provides an effective adaptive search-control mechanism for constrained maritime moving-target observation scheduling.