9 citations · 18 across the 11 of their papers we have counts for
5 papers · 1 filter
Agent Benchmarks Fail Public Sector Requirements
Jonathan Rystrøm, Chris Schmitz, Karolina Korgul +2
Deploying Large Language Model-based agents (LLM agents) in the public sector requires assuring that they meet the stringent legal, procedural, and structural requirements of publi…
Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency
Jan Batzner, Volker Stocker, Bingjun Tang +4
Synthetic personae experiments have become a prominent method in Large Language Model alignment research, yet the representativeness and ecological validity of these personae vary…
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Anka Reuel, Avijit Ghosh, Jenny Chim +32
Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general c…
Oversight Structures for Agentic AI in Public-Sector Organizations
Chris Schmitz, Jonathan Rystrøm, Jan Batzner
This paper finds that the introduction of agentic AI systems intensifies existing challenges to traditional public sector oversight mechanisms -- which rely on siloed compliance un…
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
Jan Batzner, Volker Stocker, Stefan Schmid +1
Large language models (LLMs) are increasingly shaping citizens' information ecosystems. Products incorporating LLMs, such as chatbots and AI Companions, are now widely used for dec…