Racial Disparity in Natural Language Processing: A Case Study of Social Media African-American English
arXiv:1707.00061
Abstract
We highlight an important frontier in algorithmic fairness: disparity in the quality of natural language processing algorithms when applied to language from authors of different social groups. For example, current systems sometimes analyze the language of females and minorities more poorly than they do of whites and males. We conduct an empirical analysis of racial disparity in language identification for tweets written in African-American English, and discuss implications of disparity in NLP.
Presented as a talk at the 2017 Workshop on Fairness, Accountability, and Transparency in Machine Learning (FAT/ML 2017)
Cited by in corpus (5)
- Stereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations
- Towards generalisable hate speech detection: a review on obstacles and solutions
- Examining Racial Bias in an Online Abuse Corpus with Structural Topic Modeling
- Fairkit, Fairkit, on the Wall, Who's the Fairest of Them All? Supporting Data Scientists in Training Fair Models
- Inflating Topic Relevance with Ideology: A Case Study of Political Ideology Bias in Social Topic Detection Models