Search papers, labs, and topics across Lattice.
This paper introduces OmniRSCLIP, a novel contrastive learning framework designed to integrate multi-source remote sensing data, overcoming the limitations of existing RGB-centric CLIP architectures. By employing Spectral-Spatial Basis Decomposition (SSBD), the method adapts pretrained CLIP embeddings to accommodate diverse sensor inputs without losing visual knowledge. Experimental results demonstrate that OmniRSCLIP maintains robust performance in RGB tasks while effectively extending capabilities to SAR, MSI, and HSI modalities, showcasing its versatility in remote sensing applications.
OmniRSCLIP achieves strong performance across heterogeneous remote sensing modalities while preserving the advantages of RGB-based models.
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imaging (MSI), and hyperspectral imaging (HSI). To address this limitation, we propose OmniRSCLIP, an end-to-end contrastive learning framework that supports multi-source sensor inputs for remote sensing vision-language modeling. The key idea is to extend CLIP beyond its fixed RGB input interface without breaking the pretrained visual knowledge. To this end, OmniRSCLIP introduces Spectral-Spatial Basis Decomposition (SSBD), which formulates arbitrary-channel adaptation as a basis recomposition problem: pretrained CLIP patch embeddings provide transferable spatial bases, while wavelength-conditioned coefficients span sensor-specific embedding kernels within a constrained visual prior space. This design avoids forcing heterogeneous sensors into a fixed-channel input space, while aligning them in a unified image-text semantic space. We further introduce a spectral-context-aware mask-based contrastive learning scheme to suppress modality-specific redundant features and enhance fine-grained image-text alignment. Finally, to support multi-modal training, we construct OmniRS5M, the first large-scale remote sensing image-text corpus covering RGB, SAR, MSI, and HSI. Experiments on retrieval, zero-shot classification, and semantic localization show that OmniRSCLIP preserves strong RGB-domain performance while effectively extending CLIP to heterogeneous remote sensing modalities.