Search papers, labs, and topics across Lattice.
This study uncovers a phenomenon termed "model hypnosis," where seemingly trivial cues in prompts can be combined to exert significant control over AI model behavior, affecting various model families and scales. The findings reveal that these hypnotic prompts can transfer across models, indicating a pervasive vulnerability in AI systems. This has critical implications for AI safety and interpretability, as subtle textual variations can lead to substantial shifts in model outputs.
Subtle prompt cues can exert powerful control over AI models, revealing a vulnerability that challenges our understanding of AI behavior.
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.