Search papers, labs, and topics across Lattice.
This paper critiques the prevailing view of language models as mere technical artifacts, emphasizing their dependence on the linguistic contexts provided by their training data. By examining Italian language models, the author highlights the implications of training on translated and synthetic data, questioning whether these models genuinely represent the Italian language or language in general. The study advocates for a clearer distinction between language models as technical products and those intended for linguistic exploration, suggesting that this distinction could lead to more nuanced understandings of the languages we wish to model.
Language models trained on synthetic data may not reflect the true nature of the languages they represent, challenging our assumptions about their utility in linguistic studies.
Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want language models to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between language models designed as technical products and language models designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.