Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
arXiv:1607.06520
Abstract
The blind application of machine learning runs the risk of amplifying biases present in data. Such a danger is facing us with word embedding, a popular framework to represent text data as vectors which has been used in many machine learning and natural language processing tasks. We show that even word embeddings trained on Google News articles exhibit female/male gender stereotypes to a disturbing extent. This raises concerns because their widespread use, as we describe, often tends to amplify these biases. Geometrically, gender bias is first shown to be captured by a direction in the word embedding. Second, gender neutral words are shown to be linearly separable from gender definition words in the word embedding. Using these properties, we provide a methodology for modifying an embedding to remove gender stereotypes, such as the association between between the words receptionist and female, while maintaining desired associations such as between the words queen and female. We define metrics to quantify both direct and indirect gender biases in embeddings, and develop algorithms to "debias" the embedding. Using crowd-worker evaluation as well as standard benchmarks, we empirically demonstrate that our algorithms significantly reduce gender bias in embeddings while preserving the its useful properties such as the ability to cluster related concepts and to solve analogy tasks. The resulting embeddings can be used in applications without amplifying gender bias.
Cited by in corpus (70)
- The Real-World-Weight Cross-Entropy Loss Function: Modeling the Costs of Mislabeling
- Identifying and Correcting Label Bias in Machine Learning
- Stereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations
- Auditing Search Engines for Differential Satisfaction Across Demographics
- Equalizing Gender Biases in Neural Machine Translation with Word Embeddings Techniques
- Fair k-Center Clustering for Data Summarization
- Mitigating Gender Bias in Natural Language Processing: Literature Review
- Gender Bias in Contextualized Word Embeddings
- What's in a Name? Reducing Bias in Bios without Access to Protected Attributes
- Facts as Experts: Adaptable and Interpretable Neural Memory over Symbolic Knowledge
- Evaluating the Underlying Gender Bias in Contextualized Word Embeddings
- Data, Power and Bias in Artificial Intelligence
- Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection
- Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition
- Gender-preserving Debiasing for Pre-trained Word Embeddings
- On conditional parity as a notion of non-discrimination in machine learning
- Contextualizing Hate Speech Classifiers with Post-hoc Explanation
- Measuring Bias in Contextualized Word Representations
- No computation without representation: Avoiding data and algorithm biases through diversity
- BERT has a Moral Compass: Improvements of ethical and moral values of machines
- Social Biases in NLP Models as Barriers for Persons with Disabilities
- Auditing ImageNet: Towards a Model-driven Framework for Annotating Demographic Attributes of Large-Scale Image Datasets
- REPAIR: Removing Representation Bias by Dataset Resampling
- Age and gender bias in pedestrian detection algorithms
- The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- Online Abuse toward Candidates during the UK General Election 2019: Working Paper
- Group-Fair Online Allocation in Continuous Time
- Examining Racial Bias in an Online Abuse Corpus with Structural Topic Modeling
- The Woman Worked as a Babysitter: On Biases in Language Generation
- README: REpresentation learning by fairness-Aware Disentangling MEthod
- Exposing and Correcting the Gender Bias in Image Captioning Datasets and Models
- Incorporating Priors with Feature Attribution on Text Classification
- To Transfer or Not to Transfer: Misclassification Attacks Against Transfer Learned Text Classifiers
- Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function
- Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation
- Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer
- Towards Debiasing Sentence Representations
- Mark my Word: A Sequence-to-Sequence Approach to Definition Modeling
- Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness
- Fair Division Without Disparate Impact
- Maximum Weighted Loss Discrepancy
- Dynamic Modeling and Equilibria in Fair Decision Making
- Fairness through Experimentation: Inequality in A/B testing as an approach to responsible design
- On the Morality of Artificial Intelligence
- Towards Robustifying NLI Models Against Lexical Dataset Biases
- Compass-aligned Distributional Embeddings for Studying Semantic Differences across Corpora
- Fairness and Missing Values
- Weighted Empirical Risk Minimization: Sample Selection Bias Correction based on Importance Sampling
- Towards Reducing Bias in Gender Classification
- Ethics of Artificial Intelligence in Surgery
- Towards Socially Responsible AI: Cognitive Bias-Aware Multi-Objective Learning
- MDR Cluster-Debias: A Nonlinear WordEmbedding Debiasing Pipeline
- Grammatical Gender, Neo-Whorfianism, and Word Embeddings: A Data-Driven Approach to Linguistic Relativity
- Considerations for the Interpretation of Bias Measures of Word Embeddings
- LOGAN: Local Group Bias Detection by Clustering
- Anonymized BERT: An Augmentation Approach to the Gendered Pronoun Resolution Challenge
- Fair Adversarial Networks
- Towards classification parity across cohorts
- Principled Frameworks for Evaluating Ethics in NLP Systems
- Games for Fairness and Interpretability
- Alfie: An Interactive Robot with a Moral Compass
- Going Beyond T-SNE: Exposing \texttt{whatlies} in Text Embeddings
- Defining and Evaluating Fair Natural Language Generation
- Unfairness Discovery and Prevention For Few-Shot Regression
- Word Embeddings: Stability and Semantic Change
- Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts
- Theory In, Theory Out: The uses of social theory in machine learning for social science
- RPD: A Distance Function Between Word Embeddings
- Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors
- Parallax: Visualizing and Understanding the Semantics of Embedding Spaces via Algebraic Formulae