Search papers, labs, and topics across Lattice.
To probe whether surface fluency in Arabic masks underlying morphosyntactic deficits, the authors created YallaMorph, a 600K-sample benchmark evaluating controlled morphological generation from explicit lexical and feature constraints across verbs, nouns, adjectives, and invalid configurations. Evaluating both multilingual and Arabic-specialized LLMs under diacritized and undiacritized regimes reveals that fine-grained inflectional control remains largely unsolved. Models exhibit severe performance degradation when handling cliticized structures, out-of-distribution lemmas, and morphologically rare paradigms.
Fluent surface text masks deep grammatical fragility: across 600K test cases, leading LLMs consistently break down when forced to execute fine-grained Arabic morphosyntactic control, particularly under cliticization and rare inflections.
Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.