Search papers, labs, and topics across Lattice.
This study introduces M脰VE, a comprehensive evaluation framework tailored for assessing large language models (LLMs) in the German public sector, focusing on governance dimensions such as energy consumption, provider transparency, and knowledge of German political positions. The findings indicate substantial trade-offs among models, with energy consumption varying over 60-fold and not correlating directly with model size, alongside systematic differences in information disclosure across providers. Consequently, the research underscores the necessity for public institutions to incorporate governance criteria into LLM selection, rather than relying solely on traditional performance metrics.
No single LLM excels across all governance dimensions, revealing critical trade-offs that public institutions must navigate in model selection.
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of M\"OVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.