Search papers, labs, and topics across Lattice.
This paper introduces MORFES, a benchmark specifically designed to evaluate the inflectional competence of language models in Modern Greek, consisting of 500 expert-verified items that emphasize recognition and production of inflected forms. By focusing on lower-frequency lemmas, the benchmark encourages models to demonstrate grammatical rules rather than relying on memorization. Evaluation results reveal that Sophea-Genesis-1 outperforms other models in inflectional morphology while maintaining comparable general capabilities, highlighting the need for targeted assessments in morphologically rich languages.
Sophea-Genesis-1 not only excels in inflectional morphology but also challenges the notion that larger models are inherently better at language tasks.
Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at https://huggingface.co/datasets/KIEFERSA/MORFES. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at https://huggingface.co/KIEFERSA/Sophea-Genesis-1, leads on inflectional morphology while matching similarly sized models in general capability.