collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages

Sharif Kazemi, Tanya Popli, Neil K. R. Sehgal +7

Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-bas…

cs.CL2024

HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter

Manuel Tonneau, Diyi Liu, Niyati Malhotra +4

To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in eval…

cs.CL2024

Evaluating Deduplication Techniques for Economic Research Paper Titles with a Focus on Semantic Similarity using NLP and LLMs

Doohee You, S Fraiberger

This study investigates efficient deduplication techniques for a large NLP dataset of economic research paper titles. We explore various pairing methods alongside established dista…

cs.CL2024

From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets

Manuel Tonneau, Diyi Liu, Samuel Fraiberger +3

Perceptions of hate can vary greatly across cultural contexts. Hate speech (HS) datasets, however, have traditionally been developed by language. This hides potential cultural bias…

cs.CL2024

NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data

Manuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri +4

To address the global issue of online hate, hate speech detection (HSD) systems are typically developed on datasets from the United States, thereby failing to generalize to English…