Search papers, labs, and topics across Lattice.
8
0
5
4
InfraQR reveals that infrared vision-language models can be drastically misled by structured edge-placed perturbations, with accuracy plummeting from 98.67% to 0.70%.
AirflowAttack reveals that adversarial perturbations can not only deceive infrared VLMs but also enhance their false confidence in erroneous classifications.
Infrared-aware adaptation can boost CLIP performance by over 12 points, transforming how models interpret thermal imagery.
Even with robust training techniques like EOT, a carefully crafted adversarial patch can reliably fool VIS-IR VLMs and transfer across tasks like classification, captioning, and VQA.
VLMs can be easily fooled in the real world by strategically manipulating lighting, causing them to misinterpret scenes and hallucinate nonsensical captions.
Robot control systems are shockingly vulnerable: JailWAM achieves an 84.2% success rate in jailbreaking state-of-the-art World Action Models to perform unsafe physical actions.
VLMs can be devastatingly fooled by modifying less than 2% of image pixels in a fixed, X-shaped pattern, causing them to fail spectacularly across diverse tasks like classification, captioning, and question answering.
Medical vision-language models are surprisingly brittle: clinically plausible image manipulations, like those introduced during routine acquisition and delivery, can drastically degrade their performance.