Search papers, labs, and topics across Lattice.
This study evaluates the transferability of social biases in large language models (LLMs) by analyzing 4,900 English-Swahili prompt pairs submitted to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes. The results reveal that biases do not simply transfer between languages; instead, they transform, with notable shifts in stereotype prevalence and refusal behaviors that vary significantly between English and Swahili. Specifically, while GPT-5.2 exhibited a high refusal rate in English, it showed none in Swahili, highlighting the inadequacy of English-centric bias audits for multilingual applications.
Bias in language models doesn't just transfer across languages; it transforms, revealing significant disparities in behavior and stereotype prevalence between English and Swahili.
Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini's neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, consistent with refusal behaviour anchored to English-language surface forms at the behavioural level. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.