Search papers, labs, and topics across Lattice.
This paper introduces DURA, a novel diffusion-based attack method that generates visually natural adversarial patches for Vision-Language-Action (VLA) models, effectively compromising their performance in both white-box and black-box settings. By optimizing along the latent trajectory of a pretrained diffusion model, DURA circumvents the limitations of traditional pixel-space perturbations, which often leave detectable artifacts. Experimental results demonstrate that DURA significantly outperforms existing attack methods, highlighting a critical vulnerability in VLA models that could lead to real-world safety risks.
DURA reveals that visually indistinguishable adversarial patches can exploit VLA models, posing a significant threat to their deployment in real-world robotics.
Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.