Search papers, labs, and topics across Lattice.
This paper introduces DynCur-Geo, a dynamic curiosity reward shaping framework designed to enhance the performance of UAVs in active geo-localization tasks. By adjusting the intrinsic reward based on the distance to the target, the method effectively balances exploration and goal-directed behavior, addressing the limitations of fixed reward systems that can lead to inefficient detours. Experimental results demonstrate that DynCur-Geo significantly outperforms existing baselines across various challenging scenarios, including multimodal and disaster-affected environments.
Dynamic reward shaping can dramatically improve UAV target localization by optimizing exploration strategies based on proximity to the target.
Active geo-localization enables low-altitude UAVs to search for specified targets from limited local aerial observations, supporting time-sensitive applications such as search and rescue and emergency inspection. However, multimodal target cues, restricted views, and sparse feedback make it difficult to balance exploration with target convergence. Existing curiosity-driven methods assign a fixed intrinsic-reward weight throughout search, which can continue rewarding novelty after the agent nears the target and induce detours. We propose DynCur-Geo, a dynamic curiosity framework that adjusts prediction-error intrinsic reward according to remaining target distance. A distance-aware gate encourages early exploration and shifts the policy toward goal-directed behavior near the target, while potential-based reward shaping supplies dense progress guidance. Experiments across multimodal, cross-scene, disaster-affected, and long-range settings show consistent gains over active geo-localization baselines.