Search papers, labs, and topics across Lattice.
This study investigates how scenario-wrapped prompts can exploit vulnerabilities in safety-aligned large language models (LLMs) by activating internal directions that lower refusal scores. By introducing the \textsc{Concept2Scenario} framework, the authors attribute refusal suppression to specific concepts and identify effective scenario combinations that enhance attack success rates. The findings reveal that these vulnerabilities are not only model-specific but also transferable across different LLMs, improving attack efficacy by up to 18.2 percentage points.
Scenario-wrapped prompts can significantly weaken LLM refusal safeguards, revealing shared vulnerabilities across model families that enhance attack success rates.
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relationship between the two representations. In this work, we show that scenario-wrapped prompts activate internal scenario directions whose causal steering consistently reduces refusal scores. Building on this finding, we propose \textsc{Concept2Scenario}, a concept-based attribution framework for vulnerable scenario discovery. It instantiates a broad concept space with a sparse autoencoder, attributes refusal suppression to individual concepts, translates the identified concepts into interpretable natural-language scenarios, and identifies synergistic scenario combinations through interaction attribution. Across three open-source models, two safety benchmarks, and six black-box jailbreak methods, the discovered scenarios serve as reusable priors that improve average attack success rates by up to $18.2$ percentage points. They also transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, suggesting that some scenario-level refusal vulnerabilities are shared across model families. Moreover, the identified combinations outperform their individual constituents and enable iterative attacks to succeed in fewer turns.