Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them
arXiv:1903.03862
Abstract
Word embeddings are widely used in NLP for a vast range of tasks. It was shown that word embeddings derived from text corpora reflect gender biases in society. This phenomenon is pervasive and consistent across different word embedding models, causing serious concern. Several recent works tackle this problem, and propose methods for significantly reducing this gender bias in word embeddings, demonstrating convincing results. However, we argue that this removal is superficial. While the bias is indeed substantially reduced according to the provided bias definition, the actual effect is mostly hiding the bias, not removing it. The gender bias information is still reflected in the distances between "gender-neutralized" words in the debiased embeddings, and can be recovered from them. We present a series of experiments to support this claim, for two debiasing methods. We conclude that existing bias removal techniques are insufficient, and should not be trusted for providing gender-neutral modeling.
Accepted to NAACL 2019
References in corpus (2)
Cited by in corpus (36)
- Language Models are Few-Shot Learners
- Higher-Order Explanations of Graph Neural Networks via Relevant Walks
- The Forgotten Margins of AI Ethics
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- Machine Learning on Graphs: A Model and Comprehensive Taxonomy
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- Unmasking Contextual Stereotypes: Measuring and Mitigating BERT's Gender Bias
- FrameAxis: Characterizing Microframe Bias and Intensity with Word Embedding
- Large image datasets: A pyrrhic win for computer vision?
- Hurtful Words: Quantifying Biases in Clinical Contextual Word Embeddings
- Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
- Unlearn Dataset Bias in Natural Language Inference by Fitting the Residual
- Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- A Comprehensive Analysis of Static Word Embeddings for Turkish
- Evaluating Bias In Dutch Word Embeddings
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- Censorship of Online Encyclopedias: Implications for NLP Models
- MT-Adapted Datasheets for Datasets: Template and Repository
- Online Abuse of UK MPs from 2015 to 2019: Working Paper
- Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
- Marked Attribute Bias in Natural Language Inference
- Evaluating Gender Bias in Natural Language Inference
- Understanding Undesirable Word Embedding Associations
- A Causal Inference Method for Reducing Gender Bias in Word Embedding Relations
- Race and Religion in Online Abuse towards UK Politicians: Working Paper
- Fairness in Missing Data Imputation
- Evaluating Metrics for Bias in Word Embeddings
- Discovering and Interpreting Biased Concepts in Online Communities
- Assessing Demographic Bias in Named Entity Recognition
- Towards classification parity across cohorts
- Using Adversarial Debiasing to Remove Bias from Word Embeddings
- [RE] Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation
- Debiasing Sentence Embedders through Contrastive Word Pairs
- Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space
- Second Order WinoBias (SoWinoBias) Test Set for Latent Gender Bias Detection in Coreference Resolution