Search papers, labs, and topics across Lattice.
This paper introduces Self-Evolutionary CLIP (SE-CLIP), a semi-supervised framework that enhances the adaptation of vision-language models for satellite image classification by leveraging recursive label mining. By implementing a dual-phase pipeline that begins with a few annotated samples and evolves through iterative high-confidence sample discovery, SE-CLIP effectively addresses the challenges posed by limited annotated data. The results demonstrate that SE-CLIP significantly outperforms existing semi-supervised methods on the UCM and NWPU benchmarks, showcasing its potential for practical applications in remote sensing with reduced human effort.
SE-CLIP outperforms traditional semi-supervised methods by effectively mining high-confidence samples, transforming how we adapt vision-language models to satellite imagery.
Vision-language models like CLIP have shown sig- nificant potential in handling natural images, yet their perfor- mance is often limited by the distinct characteristics of satellite imagery. While parameter-efficient adaptation techniques exist, their efficacy is frequently limited by the scarcity of annotated samples. In this letter, we propose Self-Evolutionary CLIP (SE- CLIP), a semi-supervised framework designed for recursive label mining in scene classification. The approach follows a dual-phase pipeline, where an initial warm-up on a few annotated seeds is followed by a recursive discovery phase that iteratively identifies high-confidence samples from unlabeled pools. To maintain the integrity of the evolving support set, we employ a class-balanced selection strategy that prevents the model from being dominated by easily learned categories. Results on the UCM and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches. The framework provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.