Search papers, labs, and topics across Lattice.
This study investigates the jailbreak capabilities of audio-capable foundation models by manipulating speech delivery while keeping transcript content constant. Using the PJ-Break evaluation protocol and the AdvAudio-Prosody benchmark, the authors reveal that variations in prosody, such as arousal and speaking rate, significantly enhance the effectiveness of jailbreak attempts, with some presets achieving success rates of over 40%. The findings underscore the importance of considering prosodic features in the safety evaluation of audio LLMs, as emotional delivery proves to be far more impactful than textual emotion alone.
Manipulating prosody in audio LLMs can increase jailbreak success rates by over 40%, revealing a critical vulnerability in their safety mechanisms.
Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic (38/95), Anger (35/95), and Fast (32/95) presets are all well above Neutral (4/95). The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation (27/95) on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95). Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation