Search papers, labs, and topics across Lattice.
This paper introduces OGR-MARL, an option-guided residual multi-agent reinforcement learning framework designed for heterogeneous unmanned surface vehicle (USV) cooperative pursuit in constrained port waterways. By integrating shared evader beliefs, role-conditioned option targets, adaptive rule penalties, and residual policy learning, OGR-MARL enables various MARL algorithms to effectively learn corrective actions while adhering to navigation and traffic constraints. Experimental results demonstrate that the OGR-MASAC variant achieves a 75.0% capture rate and shows strong generalization capabilities in complex port scenarios without retraining.
OGR-MARL achieves a remarkable 75% capture rate in constrained port environments, showcasing its effectiveness in heterogeneous USV coordination.
Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.