Semantics derived automatically from language corpora contain human-like biases
arXiv:1608.07187 · doi:10.1126/science.aal4230
Abstract
Artificial intelligence and machine learning are in a period of astounding growth. However, there are concerns that these technologies may be used, either with or without intention, to perpetuate the prejudice and unfairness that unfortunately characterizes many human institutions. Here we show for the first time that human-like semantic biases result from the application of standard machine learning to ordinary language---the same sort of language humans are exposed to every day. We replicate a spectrum of standard human biases as exposed by the Implicit Association Test and other well-known psychological studies. We replicate these using a widely used, purely statistical machine-learning model---namely, the GloVe word embedding---trained on a corpus of text from the Web. Our results indicate that language itself contains recoverable and accurate imprints of our historic biases, whether these are morally neutral as towards insects or flowers, problematic as towards race or gender, or even simply veridical, reflecting the {\em status quo} for the distribution of gender with respect to careers or first names. These regularities are captured by machine learning along with the rest of semantics. In addition to our empirical findings concerning language, we also contribute new methods for evaluating bias in text, the Word Embedding Association Test (WEAT) and the Word Embedding Factual Association Test (WEFAT). Our results have implications not only for AI and machine learning, but also for the fields of psychology, sociology, and human ethics, since they raise the possibility that mere exposure to everyday language can account for the biases we replicate here.
14 pages, 3 figures
References in corpus (2)
Cited by in corpus (277)
- The Kinetics Human Action Video Dataset
- Generative AI
- Universal Sentence Encoder
- Word Embeddings Quantify 100 Years of Gender and Ethnic Stereotypes
- Fairness And Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, And Mitigation Strategies
- Out of One, Many: Using Language Models to Simulate Human Samples
- What Do We Want From Explainable Artificial Intelligence (XAI)? -- A Stakeholder Perspective on XAI and a Conceptual Model Guiding Interdisciplinary XAI Research
- Fairness in Machine Learning: A Survey
- The Geometry of Culture: Analyzing Meaning through Word Embeddings
- Society-in-the-Loop: Programming the Algorithmic Social Contract
- Machine Learning and Deep Learning -- A review for Ecologists
- A Study of Generative Large Language Model for Medical Research and Healthcare
- CogView: Mastering Text-to-Image Generation via Transformers
- Gender bias and stereotypes in Large Language Models
- Fairness-Aware Ranking in Search & Recommendation Systems with Application to LinkedIn Talent Search
- Co-Writing with Opinionated Language Models Affects Users' Views
- Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them
- GenAI Against Humanity: Nefarious Applications of Generative Artificial Intelligence and Large Language Models
- Distributional Semantics and Linguistic Theory
- Explainable Artificial Intelligence: Precepts, Methods, and Opportunities for Research in Construction
- Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
- Fairway: A Way to Build Fair ML Software
- Deconfounded Image Captioning: A Causal Retrospect
- Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints
- Machine Culture
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- Automated Evaluation Of Psychotherapy Skills Using Speech And Language Technologies
- Measuring Depression Symptom Severity from Spoken Language and 3D Facial Expressions
- Stereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations
- The Cinderella Complex: Word Embeddings Reveal Gender Stereotypes in Movies and Books
- Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
- Cultural Cartography with Word Embeddings
- Improving Human-AI Collaboration With Descriptions of AI Behavior
- Confronting Abusive Language Online: A Survey from the Ethical and Human Rights Perspective
- Like trainer, like bot? Inheritance of bias in algorithmic content moderation
- Addressing Bias in Generative AI: Challenges and Research Opportunities in Information Management
- Gender Bias in Neural Natural Language Processing
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- Improving the Accuracy of Pre-trained Word Embeddings for Sentiment Analysis
- Data and its (dis)contents: A survey of dataset development and use in machine learning research
- Focus Group on Artificial Intelligence for Health
- Quantifying Bias in Automatic Speech Recognition
- Avoiding bias when inferring race using name-based approaches
- Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types
- Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks
- A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
- Unmasking Contextual Stereotypes: Measuring and Mitigating BERT's Gender Bias
- Computer says 'no': Exploring systemic bias in ChatGPT using an audit approach
- A Framework for the Computational Linguistic Analysis of Dehumanization
- Debiasing Methods for Fairer Neural Models in Vision and Language Research: A Survey
- Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
- Mitigating Gender Bias in Natural Language Processing: Literature Review
- Algorithmic Injustices: Towards a Relational Ethics
- RNNs as psycholinguistic subjects: Syntactic state and grammatical dependency
- "I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation
- A Fused Large Language Model for Predicting Startup Success
- Moral Framing and Ideological Bias of News
- The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
- The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
- Gender Bias in Contextualized Word Embeddings
- Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus
- What's in a Name? Reducing Bias in Bios without Access to Protected Attributes
- Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
- AI and Blackness: Towards moving beyond bias and representation
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- Studying the Transfer of Biases from Programmers to Programs
- SODA: A Natural Language Processing Package to Extract Social Determinants of Health for Cancer Studies
- Learning Gender-Neutral Word Embeddings
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image Representations
- 'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion
- Dual Use Concerns of Generative AI and Large Language Models
- Identifying translational science through embeddings of controlled vocabularies
- Semantic and Relational Spaces in Science of Science: Deep Learning Models for Article Vectorisation
- Inseq: An Interpretability Toolkit for Sequence Generation Models
- Gender Bias in BERT -- Measuring and Analysing Biases through Sentiment Rating in a Realistic Downstream Classification Task
- Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems
- Making Fair ML Software using Trustworthy Explanation
- WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
- Artificial Concepts of Artificial Intelligence: Institutional Compliance and Resistance in AI Startups
- Two Simple Ways to Learn Individual Fairness Metrics from Data
- First Impressions: A Survey on Vision-Based Apparent Personality Trait Analysis
- When we can trust computers (and when we can't)
- Semantic projection: recovering human knowledge of multiple, distinct object features from word embeddings
- Large scale analysis of gender bias and sexism in song lyrics
- Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
- Fairness in Agreement With European Values: An Interdisciplinary Perspective on AI Regulation
- MultiBench: Multiscale Benchmarks for Multimodal Representation Learning
- Directional Bias Amplification
- Advances in Machine Learning for the Behavioral Sciences
- Towards Fairness in Visual Recognition: Effective Strategies for Bias Mitigation
- Human Languages with Greater Information Density Increase Communication Speed, but Decrease Conversation Breadth
- Explainability for fair machine learning
- Gender-preserving Debiasing for Pre-trained Word Embeddings
- Contrastive Explanations with Local Foil Trees
- Right for the Right Reason: Training Agnostic Networks
- Contextual Word Representations: A Contextual Introduction
- Quantifying social organization and political polarization in online platforms
- Fair Allocation through Selective Information Acquisition
- CERN for AI: A Theoretical Framework for Autonomous Simulation-Based Artificial Intelligence Testing and Alignment
- Measuring Bias in Contextualized Word Representations
- Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
- Fair Representation: Guaranteeing Approximate Multiple Group Fairness for Unknown Tasks
- Evolution of emotion semantics
- Identifying Bias in AI using Simulation
- Counterfactuals and Causability in Explainable Artificial Intelligence: Theory, Algorithms, and Applications
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- Gender and content bias in Large Language Models: a case study on Google Gemini 2.0 Flash Experimental
- Robustness and Reliability of Gender Bias Assessment in Word Embeddings: The Role of Base Pairs
- No computation without representation: Avoiding data and algorithm biases through diversity
- BERT has a Moral Compass: Improvements of ethical and moral values of machines
- Much Ado About Gender: Current Practices and Future Recommendations for Appropriate Gender-Aware Information Access
- Social Biases in NLP Models as Barriers for Persons with Disabilities
- Empirical Analysis of Multi-Task Learning for Reducing Model Bias in Toxic Comment Detection
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
- Artificial mental phenomena: Psychophysics as a framework to detect perception biases in AI models
- What are the biases in my word embedding?
- Measuring Model Biases in the Absence of Ground Truth
- Being Together in Place as a Catalyst for Scientific Advance
- Towards an Enhanced Understanding of Bias in Pre-trained Neural Language Models: A Survey with Special Emphasis on Affective Bias
- Stereotype and Skew: Quantifying Gender Bias in Pre-trained and Fine-tuned Language Models
- The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- How Reliable are Model Diagnostics?
- Hi, my name is Martha: Using names to measure and mitigate bias in generative dialogue models
- Evaluating the Construct Validity of Text Embeddings with Application to Survey Questions
- Investigating Failures of Automatic Translation in the Case of Unambiguous Gender
- Online Abuse toward Candidates during the UK General Election 2019: Working Paper
- Measuring Gender Bias in Word Embeddings of Gendered Languages Requires Disentangling Grammatical Gender Signals
- Examining Racial Bias in an Online Abuse Corpus with Structural Topic Modeling
- Group-Fair Online Allocation in Continuous Time
- Evaluating Bias In Dutch Word Embeddings
- Measuring Social Bias in Knowledge Graph Embeddings
- Why resampling outperforms reweighting for correcting sampling bias with stochastic gradients
- Rethinking embedding coupling in pre-trained language models
- Who Wrote this? How Smart Replies Impact Language and Agency in the Workplace
- Quantifying Gender Biases Towards Politicians on Reddit
- Adversarial Removal of Demographic Attributes from Text Data
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- Debiasing Pre-trained Contextualised Embeddings
- Censorship of Online Encyclopedias: Implications for NLP Models
- Fair-SSL: Building fair ML Software with less data
- Synthetically generated text for supervised text analysis
- Social Norm Bias: Residual Harms of Fairness-Aware Algorithms
- Mitigating Dataset Harms Requires Stewardship: Lessons from 1000 Papers
- Online Abuse of UK MPs from 2015 to 2019: Working Paper
- Mind the GAP: A Balanced Corpus of Gendered Ambiguous Pronouns
- Speciesist Language and Nonhuman Animal Bias in English Masked Language Models
- The Geometry of Information Cocoon: Analyzing the Cultural Space with Word Embedding Models
- Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer
- Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation
- Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
- Towards Debiasing Sentence Representations
- Neutralizing Gender Bias in Word Embedding with Latent Disentanglement and Counterfactual Generation
- Fairkit, Fairkit, on the Wall, Who's the Fairest of Them All? Supporting Data Scientists in Training Fair Models
- Who is Responsible When AI Fails? Mapping Causes, Entities, and Consequences of AI Privacy and Ethical Incidents
- AraWEAT: Multidimensional Analysis of Biases in Arabic Word Embeddings
- Cognitive phantoms in LLMs through the lens of latent variables
- Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
- Reinforcement Learning for Fair Dynamic Pricing
- Evaluating Gender Bias in Natural Language Inference
- Personalized Chatbot Trustworthiness Ratings
- On the Integration of LinguisticFeatures into Statistical and Neural Machine Translation
- Marked Attribute Bias in Natural Language Inference
- Examining Gender Bias in Languages with Grammatical Gender
- Navigating Automated Hiring: Perceptions, Strategy Use, and Outcomes Among Young Job Seekers
- Understanding Undesirable Word Embedding Associations
- Avoiding Resentment Via Monotonic Fairness
- Blacks is to Anger as Whites is to Joy? Understanding Latent Affective Bias in Large Pre-trained Neural Language Models
- Dictionary-based Debiasing of Pre-trained Word Embeddings
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies
- A Causal Inference Method for Reducing Gender Bias in Word Embedding Relations
- Exploring Polarization of Users Behavior on Twitter During the 2019 South American Protests
- Examining the Presence of Gender Bias in Customer Reviews Using Word Embedding
- Cogniculture: Towards a Better Human-Machine Co-evolution
- Always Lurking: Understanding and Mitigating Bias in Online Human Trafficking Detection
- Can AI with High Reasoning Ability Replicate Human-like Decision Making in Economic Experiments?
- NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
- Happy Travelers Take Big Pictures: A Psychological Study with Machine Learning and Big Data
- Towards Lexical Gender Inference: A Scalable Methodology using Online Databases
- Evaluating race and sex diversity in the world's largest companies using deep neural networks
- Identification, Interpretability, and Bayesian Word Embeddings
- Race and Religion in Online Abuse towards UK Politicians: Working Paper
- Towards Robustifying NLI Models Against Lexical Dataset Biases
- Compass-aligned Distributional Embeddings for Studying Semantic Differences across Corpora
- Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
- generAItor: Tree-in-the-Loop Text Generation for Language Model Explainability and Adaptation
- Legal and ethical implications of applications based on agreement technologies: the case of auction-based road intersections
- Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences
- The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
- On Transferability of Bias Mitigation Effects in Language Model Fine-Tuning
- Adversarial Learning for Debiasing Knowledge Graph Embeddings
- Enhancing Deliberativeness: Evaluating the Impact of Multimodal Reflection Nudges
- Exploring Stereotypes and Biased Data with the Crowd
- Retiring Adult: New Datasets for Fair Machine Learning
- Embedding-based Qualitative Analysis of Polarization in Turkey
- Social Connection Induces Cultural Contraction: Evidence from Hyperbolic Embeddings of Social and Semantic Networks
- Assessing Demographic Bias in Named Entity Recognition
- Reconsidering Requirements Engineering: Human-AI Collaboration in AI-Native Software Development
- Understanding Gender Bias in AI-Generated Product Descriptions
- Discovering and Interpreting Biased Concepts in Online Communities
- Bots against Bias: Critical Next Steps for Human-Robot Interaction
- Argument from Old Man's View: Assessing Social Bias in Argumentation
- Investigating Societal Biases in a Poetry Composition System
- Fairness Definitions in Language Models Explained
- Social Perception of Faces in a Vision-Language Model
- Identification of Bias Against People with Disabilities in Sentiment Analysis and Toxicity Detection Models
- Text-based inference of moral sentiment change
- Evaluating Metrics for Bias in Word Embeddings
- (Ir)rationality in AI: State of the Art, Research Challenges and Open Questions
- The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability
- Considerations for the Interpretation of Bias Measures of Word Embeddings
- Neural Embeddings of Scholarly Periodicals Reveal Complex Disciplinary Organizations
- A Set of Distinct Facial Traits Learned by Machines Is Not Predictive of Appearance Bias in the Wild
- History Playground: A Tool for Discovering Temporal Trends in Massive Textual Corpora
- Can We Derive Explicit and Implicit Bias from Corpus?
- MDR Cluster-Debias: A Nonlinear WordEmbedding Debiasing Pipeline
- Biased Embeddings from Wild Data: Measuring, Understanding and Removing
- Spinning Sequence-to-Sequence Models with Meta-Backdoors
- Algorithmic Fairness Datasets: the Story so Far
- Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI
- No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
- Discrete Word Embedding for Logical Natural Language Understanding
- Obstructing Classification via Projection
- Gender Bias Hidden Behind Chinese Word Embeddings: The Case of Chinese Adjectives
- A Source-Criticism Debiasing Method for GloVe Embeddings
- The presence of occupational structure in online texts based on word embedding NLP models
- From Symbols to Embeddings: A Tale of Two Representations in Computational Social Science
- Assessing the Reliability of Word Embedding Gender Bias Measures
- Statistical discrimination in learning agents
- Automatically Inferring Gender Associations from Language
- Debiasing Convolutional Neural Networks via Meta Orthogonalization
- Detecting discriminatory risk through data annotation based on Bayesian inferences
- [RE] Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation
- Fairness-Aware Online Personalization
- LOGAN: Local Group Bias Detection by Clustering
- Fair Adversarial Networks
- Second Order WinoBias (SoWinoBias) Test Set for Latent Gender Bias Detection in Coreference Resolution
- Unpacking the Interdependent Systems of Discrimination: Ableist Bias in NLP Systems through an Intersectional Lens
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models
- Adversarial Examples Generation for Reducing Implicit Gender Bias in Pre-trained Models
- Constrained Non-Affine Alignment of Embeddings
- "Amazing, They All Lean Left" -- Analyzing the Political Temperaments of Current LLMs
- Evaluation of Human and Machine Face Detection using a Novel Distinctive Human Appearance Dataset
- Semantic coordinates analysis reveals language changes in the AI field
- Hacia los Comités de Ética en Inteligencia Artificial
- OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings
- Assessing gender bias in medical and scientific masked language models with StereoSet
- Debiasing Sentence Embedders through Contrastive Word Pairs
- Towards a Flexible Framework for Algorithmic Fairness
- Scared into Action: How Partisanship and Fear are Associated with Reactions to Public Health Directives
- Categorization in the Wild: Generalizing Cognitive Models to Naturalistic Data across Languages
- A Large-Scale, Automated Study of Language Surrounding Artificial Intelligence
- How Fair is Fairness-aware Representative Ranking and Methods for Fair Ranking
- Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors
- Games for Fairness and Interpretability
- Cross-Lingual Probing and Community-Grounded Analysis of Gender Bias in Low-Resource Bengali
- Detecting Cross-Geographic Biases in Toxicity Modeling on Social Media
- Multimodal Interactive Learning of Primitive Actions
- Defining and Evaluating Fair Natural Language Generation
- From artificial to organic: Rethinking the roots of intelligence for digital health
- Gender Stereotype Reinforcement: Measuring the Gender Bias Conveyed by Ranking Algorithms
- Alfie: An Interactive Robot with a Moral Compass
- Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays
- Text Classification based on Multiple Block Convolutional Highways
- Deconfounding age effects with fair representation learning when assessing dementia
- Human Imperceptible Attacks and Applications to Improve Fairness
- Gender and Race Bias in Consumer Product Recommendations by Large Language Models
- Contrastive Clustering: Toward Unsupervised Bias Reduction for Emotion and Sentiment Classification
- Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts
- Sexism in the Judiciary
- Auditing for Diversity using Representative Examples
- Debiasing Methods in Natural Language Understanding Make Bias More Accessible
- StackingNet: Collective Inference Across Independent AI Foundation Models
- Collecting a Large-Scale Gender Bias Dataset for Coreference Resolution and Machine Translation
- Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?
- On the Learnability of Programming Language Semantics
- Evaluating Gender Bias in Hindi-English Machine Translation
- Using Sociolinguistic Variables to Reveal Changing Attitudes Towards Sexuality and Gender