3 papers
cs.CY2026
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
Manuel Tonneau, Dylan Thurgood, Diyi Liu +7
Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at…
cs.CL2026
Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias
Manuel Tonneau, Neil K. R. Seghal, Niyati Malhotra +7
Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups. This ap…
cs.CL2025
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
Manuel Tonneau, Diyi Liu, Niyati Malhotra +4
To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in eval…