Search papers, labs, and topics across Lattice.
The M\"OVE benchmark evaluates 39 large language models (LLMs) specifically for the German public sector, addressing the limitations of existing benchmarks that are often English-centric and focused solely on task performance. It incorporates both performance criteria, such as summarization and question answering, and governance criteria, including hallucination tendencies and alignment with German constitutional values. The findings reveal that no single model excels across all evaluation criteria, highlighting the complexity of model selection in this context and the inadequacy of size as a predictor of performance quality.
No single LLM outperforms others across all criteria, revealing the nuanced challenges of model selection in the German public sector.
We present M\"OVE (Modelle f\"ur die \"Offentliche Verwaltung Evaluieren), a holistic benchmark for evaluating large language models (LLMs) in the context of the German public sector. While LLMs are increasingly adopted in public administration, model selection remains largely ad hoc, and existing benchmarks offer limited guidance: they are predominantly English-centric, US-centric in content, and focus exclusively on task performance. M\"OVE addresses these gaps by evaluating 39 models across two complementary dimensions. Performance criteria cover summarization, question answering, and topic extraction. Governance criteria assess hallucination tendencies, energy consumption, provider transparency, and alignment with German constitutional values and knowledge about positions by German political parties. In total, we utilize ten German-language datasets, including gold- and silverstandard datasets that we constructed to reflect public-administration domains. We employ a multi-metric evaluation strategy combining classical NLP metrics, embedding-based methods, and LLM-as-a-judge approaches. Our results show that no single model dominates across all criteria: top performers differ between tasks, and model size alone is a poor predictor of quality. We further evaluate the benchmark itself, analyzing its statistical precision, LLM judge reliability, the impact of our private datasets on model rankings, the sensitivity of our results to prompt formulation, and the validity of our energy consumption estimates. M\"OVE is designed as a living benchmark under active development; results are publicly available at https://moeve.bundesdruckerei.de/.