Search papers, labs, and topics across Lattice.
This paper introduces GeoThreat, a novel transferable targeted adversarial attack method specifically designed for large vision-language models (LVLMs) in the context of remote sensing image interpretation. By leveraging both conceptual representations from class tokens and perceptual representations derived from patch tokens, GeoThreat effectively manipulates local visual cues to achieve desired semantic outcomes under black-box conditions. Extensive experiments reveal that GeoThreat outperforms existing methods in terms of both transferability and controllability, highlighting its potential to enhance the robustness assessment of LVLMs in remote sensing applications.
GeoThreat achieves superior transferability and controllability in adversarial attacks on LVLMs, redefining how we assess model robustness in remote sensing contexts.
Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic understanding. Existing studies mainly focus on corrupting visual inputs to induce predefined erroneous responses in general vision-language tasks, whereas corresponding investigations in remote sensing fields remain largely underexplored. Compared with natural image understanding, remote sensing image interpretation requires joint reasoning over local discriminative cues and global scene context. This poses additional challenges to achieving transferable semantic manipulation toward specified responses under black-box settings. To tackle these challenges, we propose GeoThreat, a transferable targeted adversarial attack method against LVLMs for remote sensing image interpretation. Specifically, GeoThreat modulates adversarial representations in accordance with the target content at both conceptual and perceptual levels. The class tokens from surrogate image encoders are employed as conceptual representations, while perceptual representations are distilled from patch tokens of the adversarial example through collaborative importance estimation. Beyond merely rolling out attention scores across layers, we incorporate adversarial-target similarity gradients to more faithfully characterize the relevance of local visual cues to the intended semantic manipulation. The perceptual representations are then dynamically aligned with target patch tokens in a cross-attentive manner, facilitating the adaptation of local cues toward designated semantic details. Finally, adversarial perturbations are iteratively updated via ensemble-based joint optimization of conceptual calibration and perceptual adaptation. Extensive experiments across diverse LVLMs demonstrate the superiority of GeoThreat in both transferability and controllability.