Search papers, labs, and topics across Lattice.
This study investigates whether text generated by multilingual large language models (MLLMs) exhibits characteristics of translationese, a phenomenon where translated text retains traces of its source language. By employing high-accuracy classification models and linguistic feature analyses across five languages, the authors compare MLLM outputs to both human-written and non-translated baselines. The findings reveal that MLLM-generated text does indeed exhibit translationese, but it differs significantly from text produced through direct translation, highlighting unique markers of MLLM interference.
MLLM-generated text shows distinct signs of translationese, but surprisingly, it diverges from traditional translation patterns in significant ways.
Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs' generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.