Search papers, labs, and topics across Lattice.
2
0
4
0
The Intermodal Dual-MAE Framework with a Similarity-based RL Masking Strategy (SRLM) is proposed, which adaptively masks informative positions by leveraging cross-modal similarity and reinforcement learning, thus narrowing the modality gap.
Adversarial images in CLIP reveal a consistent directional bias that can be exploited to enhance model robustness, leading to surprising cases where adversarial accuracy surpasses clean accuracy.