9 papers
BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian
Jophin John, Michael Hoffmann, Jan Fillies +2
Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introdu…
Algospeak, Hiding in the Open: The Trade-off Between Legible Meaning and Detection Avoidance
Jan Fillies, Ronald E. Robertson, Jeffrey Hancock
As large language models (LLMs) increasingly mediate both content generation and moderation, linguistic evasion strategies known as Algospeak have intensified the coevolution betwe…
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
David Hartmann, Manuel Tonneau, Angelie Kraft +7
The closure of Perspective API at the end of 2026 discards what has functioned as the de facto standard for automated toxicity measurement in NLP, CSS, and LLM evaluation research.…
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
Peiran Li, Jan Fillies, Adrian Paschke
Augmenting toxic language data in a controllable and class-specific manner is crucial for improving robustness in toxicity classification, yet remains challenging due to limited su…
Designing and Evaluating Malinowski's Lens: An AI-Native Educational Game for Ethnographic Learning
Michael Hoffmann, Jophin John, Jan Fillies +1
This study introduces 'Malinowski's Lens', the first AI-native educational game for anthropology that transforms Bronislaw Malinowski's 'Argonauts of the Western Pacific' (1922) in…
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
Jan Fillies, Michael Peter Hoffmann, Rebecca Reichel +3
A lack of demographic context in existing toxic speech datasets limits our understanding of how different age groups communicate online. In collaboration with funk, a German public…