2 papers
cs.CY2026
Agent Benchmarks Fail Public Sector Requirements
Jonathan Rystrøm, Chris Schmitz, Karolina Korgul +2
Deploying Large Language Model-based agents (LLM agents) in the public sector requires assuring that they meet the stringent legal, procedural, and structural requirements of publi…
cs.CL2025
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…