Search papers, labs, and topics across Lattice.
This paper investigates the limitations of existing semantic-shift jailbreaks in foundation models, revealing that the effectiveness of these attacks is significantly influenced by the contextual information used. By systematically analyzing the semantic-shift capabilities of various contexts, the authors introduce Iterative Context Optimization (ICO), a framework that optimizes context iteratively based on feedback from the target model. Experimental results show that ICO achieves an average attack success rate of 74.6%, outperforming eight state-of-the-art baselines and highlighting the importance of context in enhancing jailbreak effectiveness.
Contextual information can dramatically amplify the success of semantic-shift jailbreaks, with a new framework achieving a 74.6% attack success rate.
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful terms in original harmful questions with benign alternatives and leveraging contextual information to induce the target model to reinterpret these alternatives as their corresponding harmful concepts. However, existing semantic-shift jailbreaks often achieve limited effectiveness. In this work, we reveal that this limitation arises from overlooking the semantic-shift capability of contexts. Through systematic analysis, we find that contexts exhibit substantially different abilities in inducing semantic shifts: contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks. Based on this finding, we systematically identify and distill the characteristics of effective contexts and propose a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization (ICO). In each iteration, ICO leverages these characteristics and feedback from the target model to optimize contexts. Extensive experiments on three datasets and eight target foundation models demonstrate that ICO consistently outperforms eight state-of-the-art baselines, achieving an average attack success rate of 74.6%.