Search papers, labs, and topics across Lattice.
This paper investigates the phenomenon of Unsafe Induction Attacks, where adversaries manipulate benign inputs to be misclassified as unsafe by multimodal guard models, leading to false positives. The study highlights a critical gap in safety measures, revealing that such attacks can significantly degrade service availability and user trust by causing legitimate requests to be rejected. Through the introduction of Unsafe Semantic Distillation (USD), the authors demonstrate a method that aligns adversarial perturbations with unsafe content representations, achieving an 84% success rate against four leading guard models in realistic scenarios.
Adversaries can exploit multimodal guardrails to misclassify safe inputs as unsafe, achieving an alarming 84% success rate in disrupting service availability.
Multimodal guard models have emerged as critical safety components for screening content in vision-language systems. While adversarial research has extensively studied jailbreaking attacks that produce false negatives, the inverse threat of inducing false positives on benign inputs remains unexplored. We introduce Unsafe Induction Attacks, where adversaries distribute imperceptibly perturbed safe images that trigger guard models to reject legitimate user requests, causing a "Boy Who Cried Wolf" effect that degrades service availability and erodes trust. This reveals an availability failure mode in deployed safety filters. To realize this threat under diverse user prompts, we propose Unsafe Semantic Distillation (USD), which aligns adversarial perturbations with distributional representations of unsafe content rather than prompt-specific instances. Evaluated on four state-of-the-art guard models across realistic user simulation scenarios, USD achieves 84% attack success rates, outperforming existing methods and exposing fundamental vulnerabilities in current multimodal safety architectures.