Search papers, labs, and topics across Lattice.
This study investigates the effectiveness of safety alignment in large language models (LLMs) across four low-resource African languages鈥擳wi, Hausa, Amharic, and Swahili鈥攗sing a novel dataset called LoDNA that pairs literal translations with culturally relevant prompts. The researchers employ a latent geometric framework to analyze hidden-state refusal representations, revealing that harmful prompts retain less than 10% of the English refusal signal in most language-model pairs. These results indicate that the assumption of universal safety mechanisms across languages is flawed, highlighting significant vulnerabilities in multilingual safety alignment for low-resource languages.
Cross-lingual safety transfer in LLMs is a mirage, with harmful prompts losing over 90% of their safety signals in low-resource languages.
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.