Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases
arXiv:2006.03955 · doi:10.1145/3461702.3462536
Abstract
With the starting point that implicit human biases are reflected in the statistical regularities of language, it is possible to measure biases in English static word embeddings. State-of-the-art neural language models generate dynamic word embeddings dependent on the context in which the word appears. Current methods measure pre-defined social and intersectional biases that appear in particular contexts defined by sentence templates. Dispensing with templates, we introduce the Contextualized Embedding Association Test (CEAT), that can summarize the magnitude of overall bias in neural language models by incorporating a random-effects model. Experiments on social and intersectional biases show that CEAT finds evidence of all tested biases and provides comprehensive information on the variance of effect magnitudes of the same bias in different contexts. All the models trained on English corpora that we study contain biased representations. Furthermore, we develop two methods, Intersectional Bias Detection (IBD) and Emergent Intersectional Bias Detection (EIBD), to automatically identify the intersectional biases and emergent intersectional biases from static word embeddings in addition to measuring them in contextualized word embeddings. We present the first algorithmic bias detection findings on how intersectional group members are strongly associated with unique emergent biases that do not overlap with the biases of their constituent minority identities. IBD and EIBD achieve high accuracy when detecting the intersectional and emergent biases of African American females and Mexican American females. Our results indicate that biases at the intersection of race and gender associated with members of multiple minority groups, such as African American females and Mexican American females, have the highest magnitude across all neural language models.
19 pages, 2 figures, 4 tables
References in corpus (1)
Cited by in corpus (21)
- Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation
- Gender Bias in Word Embeddings: A Comprehensive Analysis of Frequency, Syntax, and Semantics
- A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
- Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks
- Debiasing Methods for Fairer Neural Models in Vision and Language Research: A Survey
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias
- Large scale analysis of gender bias and sexism in song lyrics
- AI Ethics: A Bibliometric Analysis, Critical Issues, and Key Gaps
- Mimetic Models: Ethical Implications of AI that Acts Like You
- Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
- Mapping the individual, social, and biospheric impacts of Foundation Models
- Towards an Enhanced Understanding of Bias in Pre-trained Neural Language Models: A Survey with Special Emphasis on Affective Bias
- Intersectional Inquiry, on the Ground and in the Algorithm
- Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
- Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
- Fairness Definitions in Language Models Explained
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
- No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
- Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
- Algorithmic Fairness Datasets: the Story so Far