Search papers, labs, and topics across Lattice.
This paper reveals that the perceived robustness of flow-matching vision-language-action (VLA) models against adversarial attacks is largely an illusion, as previous methods overlooked the impact of the multi-step denoising ordinary differential equation (ODE). The authors introduce DRIFT, a novel adversarial patch attack that targets the denoising velocity field at the first step, demonstrating that this approach is both more effective and efficient than broader multi-step attacks. Their experiments on pi0 and pi0.5 across four LIBERO suites show that DRIFT can compromise nearly all solvable tasks with a single patch, significantly outperforming existing action- and embedding-space attack methods.
Attacking just the first denoising step of flow-matching VLAs with a single adversarial patch can break nearly all tasks, challenging assumptions about model robustness.
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.