Search papers, labs, and topics across Lattice.
The authors investigate the safety vulnerability of full-duplex audio models to mid-generation speech interruptions via DuplexJail, an attack delivering request-independent spoken prompts over the user channel. Evaluating across four models on 720 harmful prompts, interruption timed via fixed delays or streaming refusal cues increased attack success rates by up to +39.3 percentage points, achieving roughly 49% harmful completion on PersonaPlex-RL. These results expose how temporal turn-taking dynamics create an unmonitored side-channel that degrades safety alignment in streaming multimodal architectures.
Simply cutting off a voice assistant mid-sentence with generic audio is enough to break its safety guardrails, spiking jailbreak success rates by up to 39 percentage points.
Full-duplex speech models accept user speech while generating responses, creating an underexplored attack surface. We introduce DuplexJail, which delivers fixed, request-independent spoken prompts through the user audio channel. We compare fixed-delay interruption after the harmful request ends with refusal-triggered interruption following a cue in the model's streaming text. Across four open-source models and 720 harmful requests from AdvBench and HarmBench, fixed-delay interruption raises whole-response attack success rates on AdvBench to 40.3% for PersonaPlex and 48.7% for PersonaPlex-RL, increases of +33.8 and +39.3 percentage points. The refusal-triggered policy reaches 35.6% and 48.6%, respectively, with all trials scored regardless of whether an interruption occurs. Selected conditions also increase FLM-Audio's harmful-response rate, while BayLing-Duplex shows decreases. These findings identify spoken interruption as a jailbreak attack vector and motivate evaluating safety throughout ongoing full-duplex interaction.