Search papers, labs, and topics across Lattice.
GaelEval, a new benchmark, assesses LLM performance in Scottish Gaelic across morphosyntax, translation, and cultural knowledge, revealing uneven capabilities. Evaluating 19 LLMs, Gemini 3 Pro Preview surpasses human-level accuracy (83.3% vs 78.1%) on the morphosyntactic task. Proprietary models outperform open-weight models, while Gaelic prompting provides a marginal benefit in linguistic tasks but hinders cultural knowledge Q&A.
Gemini 3 Pro Preview unexpectedly beats human experts on Scottish Gaelic grammar, highlighting the potential for LLMs to excel in low-resource languages.
Multilingual large language models (LLMs) often exhibit emergent'shadow'capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally grounded translation benchmark and (iii) a large-scale cultural knowledge Q&A task. Evaluating 19 LLMs against a fluent-speaker human baseline ($n=30$), we find that Gemini 3 Pro Preview achieves $83.3\%$ accuracy on the linguistic task, surpassing the human baseline ($78.1\%$). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+$2.4\%$). On the cultural task, leading models exceed $90\%$ accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and shows a consistent performance gap favouring proprietary over open-weight models.