Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
arXiv:1607.06520
Abstract
The blind application of machine learning runs the risk of amplifying biases present in data. Such a danger is facing us with word embedding, a popular framework to represent text data as vectors which has been used in many machine learning and natural language processing tasks. We show that even word embeddings trained on Google News articles exhibit female/male gender stereotypes to a disturbing extent. This raises concerns because their widespread use, as we describe, often tends to amplify these biases. Geometrically, gender bias is first shown to be captured by a direction in the word embedding. Second, gender neutral words are shown to be linearly separable from gender definition words in the word embedding. Using these properties, we provide a methodology for modifying an embedding to remove gender stereotypes, such as the association between between the words receptionist and female, while maintaining desired associations such as between the words queen and female. We define metrics to quantify both direct and indirect gender biases in embeddings, and develop algorithms to "debias" the embedding. Using crowd-worker evaluation as well as standard benchmarks, we empirically demonstrate that our algorithms significantly reduce gender bias in embeddings while preserving the its useful properties such as the ability to cluster related concepts and to solve analogy tasks. The resulting embeddings can be used in applications without amplifying gender bias.
Cited by in corpus (146)
- Learning Transferable Visual Models From Natural Language Supervision
- The Real-World-Weight Cross-Entropy Loss Function: Modeling the Costs of Mislabeling
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
- Representation Bias in Data: A Survey on Identification and Resolution Techniques
- Identifying and Correcting Label Bias in Machine Learning
- Stereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations
- Auditing Search Engines for Differential Satisfaction Across Demographics
- Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment
- FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders
- Unmasking Contextual Stereotypes: Measuring and Mitigating BERT's Gender Bias
- Equalizing Gender Biases in Neural Machine Translation with Word Embeddings Techniques
- Fair k-Center Clustering for Data Summarization
- Modifying Memories in Transformer Models
- Mitigating Gender Bias in Natural Language Processing: Literature Review
- Gender Bias in Contextualized Word Embeddings
- What's in a Name? Reducing Bias in Bios without Access to Protected Attributes
- Evaluating the Underlying Gender Bias in Contextualized Word Embeddings
- Facts as Experts: Adaptable and Interpretable Neural Memory over Symbolic Knowledge
- Data, Power and Bias in Artificial Intelligence
- Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis
- Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection
- Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition
- Gender-preserving Debiasing for Pre-trained Word Embeddings
- Contextualizing Hate Speech Classifiers with Post-hoc Explanation
- On conditional parity as a notion of non-discrimination in machine learning
- Measuring Bias in Contextualized Word Representations
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- BERT has a Moral Compass: Improvements of ethical and moral values of machines
- No computation without representation: Avoiding data and algorithm biases through diversity
- Social Biases in NLP Models as Barriers for Persons with Disabilities
- REPAIR: Removing Representation Bias by Dataset Resampling
- Auditing ImageNet: Towards a Model-driven Framework for Annotating Demographic Attributes of Large-Scale Image Datasets
- Learning to Ask: Conversational Product Search via Representation Learning
- Age and gender bias in pedestrian detection algorithms
- Regional Differences in Information Privacy Concerns After the Facebook-Cambridge Analytica Data Scandal
- Uncovering Bias in Personal Informatics
- The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- Stereotype and Skew: Quantifying Gender Bias in Pre-trained and Fine-tuned Language Models
- Hi, my name is Martha: Using names to measure and mitigate bias in generative dialogue models
- Online Abuse toward Candidates during the UK General Election 2019: Working Paper
- Investigating Failures of Automatic Translation in the Case of Unambiguous Gender
- Improving Gender Translation Accuracy with Filtered Self-Training
- Evaluating Bias In Dutch Word Embeddings
- Group-Fair Online Allocation in Continuous Time
- Examining Racial Bias in an Online Abuse Corpus with Structural Topic Modeling
- AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini
- Detecting Gender Bias in Transformer-based Models: A Case Study on BERT
- An Interpretability Illusion for BERT
- The Woman Worked as a Babysitter: On Biases in Language Generation
- Debiasing Pre-trained Contextualised Embeddings
- Censorship of Online Encyclopedias: Implications for NLP Models
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- Intersectional Inquiry, on the Ground and in the Algorithm
- AI Fairness via Domain Adaptation
- Disembodied Machine Learning: On the Illusion of Objectivity in NLP
- README: REpresentation learning by fairness-Aware Disentangling MEthod
- Exposing and Correcting the Gender Bias in Image Captioning Datasets and Models
- Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function
- To Transfer or Not to Transfer: Misclassification Attacks Against Transfer Learned Text Classifiers
- Fairness and Robustness in Invariant Learning: A Case Study in Toxicity Classification
- Incorporating Priors with Feature Attribution on Text Classification
- Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer
- Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation
- Towards Debiasing Sentence Representations
- Case Studies on using Natural Language Processing Techniques in Customer Relationship Management Software
- Extending Challenge Sets to Uncover Gender Bias in Machine Translation: Impact of Stereotypical Verbs and Adjectives
- Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness
- Marked Attribute Bias in Natural Language Inference
- Mark my Word: A Sequence-to-Sequence Approach to Definition Modeling
- On the Privacy Risks of Algorithmic Fairness
- AraWEAT: Multidimensional Analysis of Biases in Arabic Word Embeddings
- Fair Division Without Disparate Impact
- Causal Attention for Vision-Language Tasks
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies
- Maximum Weighted Loss Discrepancy
- Dynamic Modeling and Equilibria in Fair Decision Making
- Fairness through Experimentation: Inequality in A/B testing as an approach to responsible design
- Dictionary-based Debiasing of Pre-trained Word Embeddings
- On the Morality of Artificial Intelligence
- Fairness and Missing Values
- Towards Robustifying NLI Models Against Lexical Dataset Biases
- Compass-aligned Distributional Embeddings for Studying Semantic Differences across Corpora
- Interpretable bias mitigation for textual data: Reducing gender bias in patient notes while maintaining classification performance
- Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
- CIDER: Context sensitive sentiment analysis for short-form text
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial Learning
- Fairness in Missing Data Imputation
- Weighted Empirical Risk Minimization: Sample Selection Bias Correction based on Importance Sampling
- Towards Reducing Bias in Gender Classification
- Ethics of Artificial Intelligence in Surgery
- Removing Spurious Features can Hurt Accuracy and Affect Groups Disproportionately
- Argument from Old Man's View: Assessing Social Bias in Argumentation
- Investigating Societal Biases in a Poetry Composition System
- Automatic Fairness Testing of Neural Classifiers through Adversarial Sampling
- Probabilistic Verification of Neural Networks Against Group Fairness
- Fairness in KI-Systemen
- Does Robustness Improve Fairness? Approaching Fairness with Word Substitution Robustness Methods for Text Classification
- MDR Cluster-Debias: A Nonlinear WordEmbedding Debiasing Pipeline
- Towards Socially Responsible AI: Cognitive Bias-Aware Multi-Objective Learning
- The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability
- Gender Bias, Social Bias and Representation: 70 Years of Bollywood
- Grammatical Gender, Neo-Whorfianism, and Word Embeddings: A Data-Driven Approach to Linguistic Relativity
- RAWLSNET: Altering Bayesian Networks to Encode Rawlsian Fair Equality of Opportunity
- Considerations for the Interpretation of Bias Measures of Word Embeddings
- Intentonomy: a Dataset and Study towards Human Intent Understanding
- VERB: Visualizing and Interpreting Bias Mitigation Techniques for Word Representations
- Fair Adversarial Networks
- Debiasing Convolutional Neural Networks via Meta Orthogonalization
- Fairness-aware Outlier Ensemble
- Towards classification parity across cohorts
- Equality of opportunity in travel behavior prediction with deep neural networks and discrete choice models
- Assessing the Reliability of Word Embedding Gender Bias Measures
- Anonymized BERT: An Augmentation Approach to the Gendered Pronoun Resolution Challenge
- Principled Frameworks for Evaluating Ethics in NLP Systems
- Gender Bias Hidden Behind Chinese Word Embeddings: The Case of Chinese Adjectives
- Obstructing Classification via Projection
- LOGAN: Local Group Bias Detection by Clustering
- Knowledge Graphs Evolution and Preservation -- A Technical Report from ISWS 2019
- Beyond traditional assumptions in fair machine learning
- Identifying Best Fair Intervention
- Theory In, Theory Out: The uses of social theory in machine learning for social science
- Word Embeddings: Stability and Semantic Change
- Parallax: Visualizing and Understanding the Semantics of Embedding Spaces via Algebraic Formulae
- Defining and Evaluating Fair Natural Language Generation
- Scared into Action: How Partisanship and Fear are Associated with Reactions to Public Health Directives
- Auditing for Diversity using Representative Examples
- Mind Your Outliers! Investigating the Negative Impact of Outliers on Active Learning for Visual Question Answering
- Sexism in the Judiciary
- Games for Fairness and Interpretability
- Going Beyond T-SNE: Exposing \texttt{whatlies} in Text Embeddings
- Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors
- Alfie: An Interactive Robot with a Moral Compass
- The Challenge of Diacritics in Yoruba Embeddings
- Unfairness Discovery and Prevention For Few-Shot Regression
- Mitigating the Position Bias of Transformer Models in Passage Re-Ranking
- Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts
- Adversarial Examples Generation for Reducing Implicit Gender Bias in Pre-trained Models
- Unpacking the Interdependent Systems of Discrimination: Ableist Bias in NLP Systems through an Intersectional Lens
- RPD: A Distance Function Between Word Embeddings
- Second Order WinoBias (SoWinoBias) Test Set for Latent Gender Bias Detection in Coreference Resolution
- Identifying and Mitigating Gender Bias in Hyperbolic Word Embeddings
- A Large-Scale, Automated Study of Language Surrounding Artificial Intelligence
- Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?
- Uncovering Implicit Gender Bias in Narratives through Commonsense Inference
- NeuTral Rewriter: A Rule-Based and Neural Approach to Automatic Rewriting into Gender-Neutral Alternatives