Search papers, labs, and topics across Lattice.
This study introduces BabelSteering, an innovative activation steering method that leverages safety signals from English to enhance safety alignment in multilingual contexts for large language models (LLMs). By evaluating its effectiveness across eight languages, the authors demonstrate that BabelSteering significantly increases the refusal rate of harmful requests by an average of 11 percentage points without compromising task utility. Notably, the method also leads to a 13 percentage point increase in the refusal of pseudo-harmful prompts, showcasing its potential as a low-cost intervention for multilingual safety alignment.
BabelSteering boosts harmful request refusals across multiple languages by an average of 11 percentage points, all while maintaining task performance.
Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting with LLMs in other languages may encounter weaker safeguards despite relying on the same systems for similarly sensitive tasks. In this work, we investigate whether safety signals learned from a high-resource language, like English, can improve multilingual safety. We propose BabelSteering, an activation steering method that acts as a lightweight inference- time intervention, using refusal directions derived from English safety supervision to generalize across languages. Our evaluation includes eight languages and jointly measures refusal of harmful requests, over-refusal, and general task utility. The results show that BabelSteering increases the refusal of harmful requests across languages, with only a marginal to no reduction in task utility but with some increase in refusal of pseudo-harmful prompts. For example, for Gemma 7B, we see an average increase in the refusal of harmful prompts across languages of 11 percentage points (pp), with individual languages like Bengali seeing an increase of 17 pp, with no loss of utility on Global MMLU, while pseudo-harmful refusals increase by 13 pp on average. We also introduce a multilingual translation-and-evaluation pipeline to facilitate future work on cross-lingual safety interventions. Overall, our findings suggest that activation steering may provide a practical, low- cost mechanism for extending English-derived safety signals to other languages. Warning: this paper contains examples with unsafe content