Search papers, labs, and topics across Lattice.
This paper introduces ConlangBench, a comprehensive benchmark for evaluating and training large language models (LLMs) on 21 constructed languages (conlangs), utilizing a dataset of over 21 million conlang-English parallel sentence pairs. The research reveals that LLMs exhibit superior performance on a posteriori conlangs, which are derived from natural languages, and demonstrates that models can effectively learn from the available data for eight conlangs. These findings underscore the potential of conlangs as a valuable resource for understanding LLM language acquisition, particularly in low-resource contexts.
LLMs learn more effectively from constructed languages derived from natural languages, revealing insights into their language acquisition processes.
Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for studying language learning in large language models (LLMs), existing conlangs remain largely underexplored in LLM research. We present ConlangBench, the first large-scale benchmark for evaluating and training LLMs on 21 existing conlangs. We collect over 21M conlang-English parallel sentence pairs (including 430K pairs across the 20 non-Esperanto conlangs) and 321K vocabulary entries. In bidirectional translation experiments, we find that models perform better on a posteriori conlangs, whose vocabularies are derived from natural languages, reflecting the design characteristics of conlangs. Training on ConlangBench also shows that models can learn all eight conlangs for which sufficient parallel corpora are available, while their learning curves vary depending on how the conlangs were created. Our findings suggest that conlangs provide a unique testbed for investigating how LLMs acquire low-resource languages.