activity
20212024
most citedRe-contextualizing Fairness in NLP: The Case of India

7 citations · 17 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2024

D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation

Aida Mostafazadeh Davani, Mark Díaz, Dylan Baker +1

While human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection. Recent studies that have critically examin…

cs.CL2024

SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes

Mukul Bhutani, Kevin Robinson, Vinodkumar Prabhakaran +2

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially pro…

cs.CL20231 cited

A Taxonomy of Rater Disagreements: Surveying Challenges & Opportunities from the Perspective of Annotating Online Toxicity

Wenbo Zhang, Hangzhi Guo, Ian D Kivlichan +3

Toxicity is an increasingly common and severe issue in online spaces. Consequently, a rich line of machine learning research over the past decade has focused on computationally det…

cs.CL20233 cited

Building Socio-culturally Inclusive Stereotype Resources with Community Engagement

Sunipa Dev, Jaya Goyal, Dinesh Tewari +2

With rapid development and deployment of generative language models in global settings, there is an urgent need to also scale our measurements of harm, not just in the number and t…

cs.CL20232 cited

SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models

Akshita Jha, Aida Davani, Chandan K. Reddy +3

Stereotype benchmark datasets are crucial to detect and mitigate social stereotypes about groups of people in NLP models. However, existing datasets are limited in size and coverag…

cs.CL2023

MD3: The Multi-Dialect Dataset of Dialogues

Jacob Eisenstein, Vinodkumar Prabhakaran, Clara Rivera +2

We introduce a new dataset of conversational speech representing English from India, Nigeria, and the United States. The Multi-Dialect Dataset of Dialogues (MD3) strikes a new bala…