Search papers, labs, and topics across Lattice.
This paper introduces BLUEPRINT, a safety-evaluation framework that dissects the conversational mechanisms behind multi-turn jailbreak attacks by integrating a social-influence strategy space with a situational context module. Using Monte Carlo Tree Search, the authors optimize combinations of 18 influence factors across four-turn dialogues, achieving near-ceiling attack success rates while minimizing the average number of queries. The study reveals that model-specific vulnerabilities exist, with all models sharing a common recovery pathway that emphasizes the importance of actionable requests in mitigating harmful dialogue outcomes.
Multi-turn jailbreak attacks reveal that framing requests as actionable can significantly reduce model vulnerability, highlighting a critical leverage point in AI safety.
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier models, BLUEPRINT achieves near-ceiling ASR on major open-weight and proprietary models, while requiring the fewest average queries (2.46). The resulting trajectories further reveal model-specific vulnerability among resistant targets: each responds to distinct influence factors and strategy transitions, yet all share a common recovery pathway-shifting toward concrete, executable task framing consistently escapes hard-refusal states. Ablations confirm operational cues matter most: making requests actionable has the largest impact, gain framing is unusually potent, and some legitimacy appeals can backfire. These findings suggest robust multi-turn safety requires monitoring not only harmful content, but also how dialogue state makes unsafe requests appear concrete and locally executable.