activity
20242026
collaborators

10 papers

cs.CY2026

The Enforcement and Feasibility of Hate Speech Moderation on Twitter

Manuel Tonneau, Dylan Thurgood, Diyi Liu +7

Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at…

cs.CL2026

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

Tiancheng Hu, Joachim Baumann, Lorenzo Lupo +3

Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human b…

cs.CL2025

Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance

Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy +1

Expert persona prompting -- assigning roles such as expert in math to language models -- is widely used for task improvement. However, prior work shows mixed results on its effecti…

cs.CL2025

HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter

Manuel Tonneau, Diyi Liu, Niyati Malhotra +4

To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in eval…

cs.CL2025

From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets

Manuel Tonneau, Diyi Liu, Samuel Fraiberger +3

Perceptions of hate can vary greatly across cultural contexts. Hate speech (HS) datasets, however, have traditionally been developed by language. This hides potential cultural bias…

cs.CL2025

Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Xinpeng Wang, Chengzhi Hu, Paul Röttger +1

Training a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful…