collaborators

5 papers

cs.CL2026

LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages

Sharif Kazemi, Tanya Popli, Neil K. R. Sehgal +7

Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-bas…

cs.AI2026

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

Sunny Rai, Jinyi Kuang, Reyhan Jamalova +7

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, soc…

cs.CY2026

The Enforcement and Feasibility of Hate Speech Moderation

Manuel Tonneau, Dylan Thurgood, Diyi Liu +7

Online hate speech is associated with harms ranging from deteriorating mental health to violence, yet how consistently platforms moderate hate, and whether enforcement is feasible…

cs.CL2026

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

Manuel Tonneau, Neil K. R. Sehgal, Niyati Malhotra +7

Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups. This ap…

cs.CL2024

HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter

Manuel Tonneau, Diyi Liu, Niyati Malhotra +4

To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in eval…