Search papers, labs, and topics across Lattice.
This paper investigates how large language models (LLMs) reflect and reinforce standard language ideologies, particularly privileging Inner Circle English norms while marginalizing non-dominant varieties. Through empirical studies and discourse analysis, it reveals that AI technologies perpetuate dominant ideologies across multiple levels, including training data and public commentary, exemplified by the controversy surrounding the term "delve." The findings highlight a "standardization paradox," where AI both homogenizes English through standard forms and pluralizes it by exposing diverse Englishes, ultimately reigniting critical discussions about language legitimacy and ownership in the context of AI systems.
AI technologies are not just tools; they actively shape and police language ideologies, privileging certain Englishes while marginalizing others.
The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc. This paper examines how AI systems, their uses and discourse on them reflect, reinforce, and occasionally challenge (standard) language ideologies, which privilege Inner Circle norms and marginalize non-dominant Englishes. Drawing on evidence from empirical studies, media commentary, social media debates, and examples from AI outputs, the paper shows that AI technologies reproduce dominant language ideologies at different levels: training data, design protocols, evaluation benchmarks, user feedback and public commentary. The analysis uses the public controversy over AI-sounding language, especially the fixation on the word delve, to illustrate how speakers of English from the Global North police the English language norms of Global South English users. The paper also identifies what Christian Mair has called a"standardisation paradox": AI may homogenize English by privileging standard forms and at the same time pluralize Englishes through exposure to wide-ranging corpora and annotation work carried out by Global South users. In doing so, the paper argues that generative AI is reigniting long-standing debates in World Englishes about standardization, legitimacy, and the ownership of English, now playing out in algorithmic systems, model training, evaluation practices, and public discourse, where non-dominant Englishes are increasingly conflated with AI-generated speech. Discussing AI systems as a site where language ideologies are (re)produced, the paper argues for more inclusive design approaches that recognize the plurality of Englishes in order to address the real-world negative consequences of treating some as more legitimate than others.