Search papers, labs, and topics across Lattice.
This study explores the use of large language models (LLMs) for predicting typological features in multilingual NLP, addressing the limitations of existing methods that lack interpretability and robustness across different resource levels. By employing an in-context learning approach with linguistic data from URIEL+ and Glottolog, the authors demonstrate that incorporating phylogenetic and geographic neighbor evidence significantly enhances prediction accuracy, particularly for low-resource languages. The findings reveal that LLMs not only outperform baseline models but also provide rationales that align with the evidence, contributing to the explainability of typological feature predictions.
LLMs can accurately predict typological features while providing interpretable rationales, even for low-resource languages, when given the right contextual evidence.
Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels and feature types remains underexplored. Given LLMs'abilities in meta-linguistic reasoning and in providing rationales, we investigate LLMs'performance in typological feature prediction via an in-context learning approach with linguistic data from URIEL+ and Glottolog. We find that zero-shot prompting is insufficient, but when given phylogenetic and geographic neighbour evidence, LLMs substantially outperform all baselines without disadvantaging low-resource languages. We further find that most LLM rationales are consistent with the provided evidence, offering a step toward explainable typological feature prediction.