34 citations · 99 across the 59 of their papers we have counts for
58 papers · 1 filter
PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models
Gianluca Barmina, Federico Torrielli, Sven Harms +7
Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpf…
Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails
Marco Antonio Stranisci, A Pranav, Rossana Damiano +2
Modern language models rely on pretraining filters to remove undesirable content from training corpora and inference-time guardrails to suppress undesirable outputs during deployme…
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
Fabian Mewes, Anne Lauscher, Vagrant Gautam
Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about reference. More recently, the interp…
Greater accessibility can amplify discrimination in generative AI
Carolin Holtermann, Minh Duc Bui, Kaitlyn Zhou +3
Hundreds of millions of people rely on large language models (LLMs) for education, work, and even healthcare. Yet these models are known to reproduce and amplify social biases pres…
Reviewing the Reviewer: LLM-Assisted Reviewer Feedback Generation for Guideline Compliance
Sukannya Purkayastha, Qile Wan, Anne Lauscher +2
Peer review is central to scientific quality, yet reliance on simple heuristics, namely lazy thinking and non-specific critiques, has threatened review quality. Prior work frames l…
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation
Carolin Holtermann, Florian Schneider, Anne Lauscher
Text-to-image (T2I) models are increasingly employed by users worldwide. However, prior research has pointed to the high sensitivity of T2I towards particular input languages - whe…