Search papers, labs, and topics across Lattice.
This study reveals a critical vulnerability in large language models related to non-imperative syntactic forms, which can be exploited to bypass safety alignment mechanisms. By evaluating 16 models with up to 70 billion parameters, the authors employ causal mediation analysis to show that the models' refusal to generate harmful content is influenced by upstream syntactic features. The research highlights that the linguistic biases present in post-training data contribute to this issue, and increasing syntactic diversity can help mitigate the risks associated with these vulnerabilities.
Non-imperative syntactic structures can undermine safety alignment in large language models, exposing them to sophisticated jailbreaks.
Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.