3 papers
cs.AI2026
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Jason Branford, Angelie Kraft
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmar…
cs.CL2026
Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure
David Hartmann, Manuel Tonneau, Angelie Kraft +7
Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers' dependence on a tool they did not control. Drawing on t…
cs.CL2025
Social Bias in Popular Question-Answering Benchmarks
Angelie Kraft, Judith Simon, Sonja Schimmler
Question-answering (QA) and reading comprehension (RC) benchmarks are commonly used for assessing the capabilities of large language models (LLMs) to retrieve and reproduce knowled…