Search papers, labs, and topics across Lattice.
China University of Petroleum-Beijing at Karamay
3
0
4
InfraQR reveals that infrared vision-language models can be drastically misled by structured edge-placed perturbations, with accuracy plummeting from 98.67% to 0.70%.
Infrared data, often overlooked, can dramatically enhance vision-language models, as shown by FusionRS's ability to improve dual-modal understanding and captioning performance.
Even with robust training techniques like EOT, a carefully crafted adversarial patch can reliably fool VIS-IR VLMs and transfer across tasks like classification, captioning, and VQA.