Search papers, labs, and topics across Lattice.
5
0
4
0
AirflowAttack reveals that adversarial perturbations can not only deceive infrared VLMs but also enhance their false confidence in erroneous classifications.
Infrared-aware adaptation can boost CLIP performance by over 12 points, transforming how models interpret thermal imagery.
Infrared data, often overlooked, can dramatically enhance vision-language models, as shown by FusionRS's ability to improve dual-modal understanding and captioning performance.
VLMs can be easily fooled in the real world by strategically manipulating lighting, causing them to misinterpret scenes and hallucinate nonsensical captions.
VLMs can be devastatingly fooled by modifying less than 2% of image pixels in a fixed, X-shaped pattern, causing them to fail spectacularly across diverse tasks like classification, captioning, and question answering.