papers

Publications (15)

cs.LG2026

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

David Hartmann, Lena Pohlmann, Lelia Hanslik +3

Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from re…

cs.CL2026

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation

David Hartmann, Manuel Tonneau, Angelie Kraft +7

The closure of Perspective API at the end of 2026 discards what has functioned as the de facto standard for automated toxicity measurement in NLP, CSS, and LLM evaluation research.…

cs.HC2026

Human-Centred LLM Privacy Audits: Findings and Frictions

Dimitri Staufer, Kirsten Morehouse, David Hartmann +1

Large language models (LLMs) learn statistical associations from massive training corpora and user interactions, and deployed systems can surface or infer information about individ…

cs.CR2026

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

Devina Jain, David Hartmann, Chuan Li

LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools colle…

cs.HC2025

Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations

David Hartmann, Amin Oueslati, Dimitri Staufer +3

Commercial content moderation APIs are marketed as scalable solutions to combat online hate speech. However, the reliance on these APIs risks both silencing legitimate speech, call…

cs.CY2025

Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society

David Hartmann, José Renato Laranjeira de Pereira, Chiara Streitbörger +1

The European legislature has proposed the Digital Services Act (DSA) and Artificial Intelligence Act (AIA) to regulate platforms and Artificial Intelligence (AI) products. We revie…