Search papers, labs, and topics across Lattice.
This paper introduces CoNav-UAV, a novel approach for target-oriented vision-and-language navigation (VLN) using a cooperative dual-altitude strategy modeled as a Stackelberg game between a high-altitude leader and a low-altitude follower. By employing Iterative Stackelberg Learning, the leader enhances its vision-language reasoning through memory-based in-context learning, while the follower improves its motion control via expert distillation, leading to a significant performance boost. Experimental results demonstrate that CoNav-UAV achieves up to a 30.8-point increase in success rate on the learning scene and outperforms existing baselines with reduced adaptation data requirements.
A cooperative dual-altitude navigation strategy can improve aerial target localization success rates by over 30% while using significantly less adaptation data.
Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locate targets given only a concise description of their appearance and surroundings. This requires global exploration and grounding as well as collision-free close-range approach, two interleaved processes difficult to reconcile within a single agent. Most existing methods transfer the ground VLN paradigm to a low-altitude UAV and compensate for its inefficient exploration with external assistance. A recent attempt deploys two UAVs at complementary altitudes yet still relies on privileged information and trains its two agents independently, precluding any mutual adaptation essential for cooperation. Here we propose CoNav-UAV, which explicitly models the task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone. To solve this game, we introduce Iterative Stackelberg Learning. The leader's high-level vision-language reasoning is refined via memory-based in-context learning, while the follower's precise motion control is updated via DAgger-style expert distillation. The alternation drives both agents toward a Stackelberg equilibrium. CoNav-UAV consistently outperforms single- and dual-agent baselines across three high-fidelity urban scenes from the AerialVLN benchmark. Success rate improves by up to 30.8 points on the learning scene, and 9.0 points under cross-scene transfer while using about 3x less adaptation data. Further analyses validate the complementary gains of the leader and follower updates and reveal robust gains yet distinct learning dynamics across VLM backbones.