Search papers, labs, and topics across Lattice.
This paper introduces DARAD, a dual-adapter and ranking-aware distillation framework designed to enhance continual remote sensing image-text retrieval (RS-ITR) amidst challenges posed by scale variation and distribution shifts. By employing a spatial fusion adapter in the visual branch and multi-expert semantic routing in the textual branch, DARAD effectively integrates new visual and textual concepts while preserving historical cross-modal ranking structures. Experimental results demonstrate that DARAD significantly outperforms existing continual learning methods, improving both adaptation to new data and retention of historical retrieval effectiveness.
DARAD not only adapts to new remote sensing data but also preserves the integrity of historical retrieval performance, a dual capability that sets a new standard in continual learning.
With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challenging because scale variation and distribution shifts in RS aggravate cross-modal alignment space distortion, making it difficult for existing continual learning (CL) methods to support reliable continual retrieval. To address this challenge, we propose DARAD, a dual-adapter and ranking-aware distillation framework that preserves the historical cross-modal ranking structure while learning new visual and textual concepts from evolving archives. Specifically, the visual branch introduces a spatial fusion adapter, which integrates coarse regional cues and fine-grained patch cues to accommodate RS scale variation while anchoring visual updates to the pretrained alignment space. The textual branch employs multi-expert semantic routing, which separates shared textual semantics from semantically specialized residuals to absorb newly emerging descriptions while constraining global text embedding drift. Furthermore, bidirectional ranking distillation uses a frozen teacher model and historical anchors to preserve the historical cross-modal ranking structure, thereby mitigating alignment space distortion across continual stages. Experiments under a multi-stage continual retrieval protocol show that DARAD achieves superior performance over existing CL methods, improving adaptation to newly arrived data while maintaining effectiveness on historical data.