Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector
Camilla Dalerci, Thilo Michael, Robin Schaefer +1
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect Engli…
cs.CL2026
MÖVE: A Holistic LLM Benchmark for the German Public Sector
Camilla Dalerci, Thilo Michael, Robin Schaefer +1
We present MÖVE (Modelle für die Öffentliche Verwaltung Evaluieren), a holistic benchmark for evaluating large language models (LLMs) in the context of the German public sector. Wh…