Search papers, labs, and topics across Lattice.
China University of Petroleum-Beijing at Karamay, Shenzhen Research Institute of Big Data
6
0
4
InfraQR reveals that infrared vision-language models can be drastically misled by structured edge-placed perturbations, with accuracy plummeting from 98.67% to 0.70%.
AirflowAttack reveals that adversarial perturbations can not only deceive infrared VLMs but also enhance their false confidence in erroneous classifications.
Infrared-aware adaptation can boost CLIP performance by over 12 points, transforming how models interpret thermal imagery.
Infrared data, often overlooked, can dramatically enhance vision-language models, as shown by FusionRS's ability to improve dual-modal understanding and captioning performance.
Even with robust training techniques like EOT, a carefully crafted adversarial patch can reliably fool VIS-IR VLMs and transfer across tasks like classification, captioning, and VQA.
VLMs can be devastatingly fooled by modifying less than 2% of image pixels in a fixed, X-shaped pattern, causing them to fail spectacularly across diverse tasks like classification, captioning, and question answering.