8 papers
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
Shaghayegh Kolli, Sina Emami, Moreno D'Incà +4
Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations c…
PrivBench: A Holistic and Modular Benchmarking Platform for Evaluating Text-to-Text Privatization
Stephen Meisenbacher, Andreea-Elena Bodea, Ahmet Bilal Akın +3
Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text privatization, where the goal…
Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy
Stephen Meisenbacher, Vlad Garbuz, Chirill Donos +5
Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and robust automatic hate speech…
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
Shaghayegh Kolli, Timo Cavelius, Nafiseh Nikeghbal +2
Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential settings, yet the visual cues that shape how these models judge people…
Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli +1
Standard accuracy benchmarks evaluate whether large language models (LLMs) reach correct answers. However, they do not test whether models maintain that answer when challenged by a…
Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models
Zhiwen You, Nafiseh Nikeghbal, Jana Diesner
Language models (LMs) can produce gendered language and stereotypes even when given neutral prompts. Most prior work on gender bias in LMs primarily examines gender through a binar…