Word Embeddings Quantify 100 Years of Gender and Ethnic Stereotypes
arXiv:1711.08412 · doi:10.1073/pnas.1720347115
Abstract
Word embeddings use vectors to represent words such that the geometry between vectors captures semantic relationship between the words. In this paper, we develop a framework to demonstrate how the temporal dynamics of the embedding can be leveraged to quantify changes in stereotypes and attitudes toward women and ethnic minorities in the 20th and 21st centuries in the United States. We integrate word embeddings trained on 100 years of text data with the U.S. Census to show that changes in the embedding track closely with demographic and occupation shifts over time. The embedding captures global social shifts -- e.g., the women's movement in the 1960s and Asian immigration into the U.S -- and also illuminates how specific adjectives and occupations became more closely associated with certain populations over time. Our framework for temporal analysis of word embedding opens up a powerful new intersection between machine learning and quantitative social science.
References in corpus (2)
Cited by in corpus (133)
- Out of One, Many: Using Language Models to Simulate Human Samples
- The Geometry of Culture: Analyzing Meaning through Word Embeddings
- Gender bias and stereotypes in Large Language Models
- Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale
- Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
- Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning
- Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them
- Evaluating Large Language Models in Theory of Mind Tasks
- Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
- A Survey of Word Embeddings Evaluation Methods
- Machine Culture
- Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases
- Measuring Depression Symptom Severity from Spoken Language and 3D Facial Expressions
- Defining and Detecting Toxicity on Social Media: Context and Knowledge are Key
- The Cinderella Complex: Word Embeddings Reveal Gender Stereotypes in Movies and Books
- Cultural Cartography with Word Embeddings
- A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle
- Gender Bias in Word Embeddings: A Comprehensive Analysis of Frequency, Syntax, and Semantics
- Principled approach to the selection of the embedding dimension of networks
- Data and its (dis)contents: A survey of dataset development and use in machine learning research
- Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment
- Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types
- FrameAxis: Characterizing Microframe Bias and Intensity with Word Embedding
- Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure
- Characterising Bias in Compressed Models
- Integrating topic modeling and word embedding to characterize violent deaths
- A Framework for the Computational Linguistic Analysis of Dehumanization
- Mitigating Gender Bias in Natural Language Processing: Literature Review
- Factors Influencing the Surprising Instability of Word Embeddings
- A Fused Large Language Model for Predicting Startup Success
- Left, Right, and Gender: Exploring Interaction Traces to Mitigate Human Biases
- Socially Responsible AI Algorithms: Issues, Purposes, and Challenges
- Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
- Learning Gender-Neutral Word Embeddings
- Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
- Hurtful Words: Quantifying Biases in Clinical Contextual Word Embeddings
- Semantic and Relational Spaces in Science of Science: Deep Learning Models for Article Vectorisation
- Large scale analysis of gender bias and sexism in song lyrics
- Whither Bias Goes, I Will Go: An Integrative, Systematic Review of Algorithmic Bias Mitigation
- Unsupervised embedding of trajectories captures the latent structure of scientific migration
- Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection
- Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis
- Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition
- Fair Allocation through Selective Information Acquisition
- Portrayal: Leveraging NLP and Visualization for Analyzing Fictional Characters
- Quantifying social organization and political polarization in online platforms
- Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
- Measuring Bias in Contextualized Word Representations
- Counterfactuals and Causability in Explainable Artificial Intelligence: Theory, Algorithms, and Applications
- No computation without representation: Avoiding data and algorithm biases through diversity
- DramatVis Personae: Visual Text Analytics for Identifying Social Biases in Creative Writing
- Empirical Analysis of Multi-Task Learning for Reducing Model Bias in Toxic Comment Detection
- A Method to Analyze Multiple Social Identities in Twitter Bios
- Social Biases in NLP Models as Barriers for Persons with Disabilities
- What are the biases in my word embedding?
- Multiparameter estimation via an ensemble of spinor atoms
- Semantic Journeys: Quantifying Change in Emoji Meaning from 2012-2018
- The citation disadvantage of clinical research
- ESR: Ethics and Society Review of Artificial Intelligence Research
- Being Together in Place as a Catalyst for Scientific Advance
- Towards an Enhanced Understanding of Bias in Pre-trained Neural Language Models: A Survey with Special Emphasis on Affective Bias
- Evaluating the Construct Validity of Text Embeddings with Application to Survey Questions
- Investigating Failures of Automatic Translation in the Case of Unambiguous Gender
- Measuring Gender Bias in Word Embeddings of Gendered Languages Requires Disentangling Grammatical Gender Signals
- Measuring Social Bias in Knowledge Graph Embeddings
- Examining Racial Bias in an Online Abuse Corpus with Structural Topic Modeling
- Evaluating Bias In Dutch Word Embeddings
- Subverting Fair Image Search with Generative Adversarial Perturbations
- A Comprehensive Analysis of Static Word Embeddings for Turkish
- Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology
- Quantifying Gender Biases Towards Politicians on Reddit
- Censorship of Online Encyclopedias: Implications for NLP Models
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- Social Norm Bias: Residual Harms of Fairness-Aware Algorithms
- The Geometry of Information Cocoon: Analyzing the Cultural Space with Word Embedding Models
- Speciesist Language and Nonhuman Animal Bias in English Masked Language Models
- Revealing Neural Network Bias to Non-Experts Through Interactive Counterfactual Examples
- Unsupervised Domain Adaptation of Contextualized Embeddings for Sequence Labeling
- Neutralizing Gender Bias in Word Embedding with Latent Disentanglement and Counterfactual Generation
- On the Integration of LinguisticFeatures into Statistical and Neural Machine Translation
- Extending Challenge Sets to Uncover Gender Bias in Machine Translation: Impact of Stereotypical Verbs and Adjectives
- Examining the Presence of Gender Bias in Customer Reviews Using Word Embedding
- Mobile Phone Data for Children on the Move: Challenges and Opportunities
- Exploring Polarization of Users Behavior on Twitter During the 2019 South American Protests
- Blacks is to Anger as Whites is to Joy? Understanding Latent Affective Bias in Large Pre-trained Neural Language Models
- Are Chess Discussions Racist? An Adversarial Hate Speech Data Set
- Hate Speech Classifiers Learn Human-Like Social Stereotypes
- Compass-aligned Distributional Embeddings for Studying Semantic Differences across Corpora
- Social Centralization and Semantic Collapse: Hyperbolic Embeddings of Networks and Text
- Identification, Interpretability, and Bayesian Word Embeddings
- Social Perception of Faces in a Vision-Language Model
- Discovering and Interpreting Biased Concepts in Online Communities
- Embedding-based Qualitative Analysis of Polarization in Turkey
- Argument from Old Man's View: Assessing Social Bias in Argumentation
- Simplicity Bias Leads to Amplified Performance Disparities
- Text-based inference of moral sentiment change
- Intentonomy: a Dataset and Study towards Human Intent Understanding
- Labor Space: A Unifying Representation of the Labor Market via Large Language Models
- Towards Fairness in Classifying Medical Conversations into SOAP Sections
- Considerations for the Interpretation of Bias Measures of Word Embeddings
- Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
- Bayesian nonparametric temporal dynamic clustering via autoregressive Dirichlet priors
- Neural Embeddings of Scholarly Periodicals Reveal Complex Disciplinary Organizations
- AI and Holistic Review: Informing Human Reading in College Admissions
- Gender Bias, Social Bias and Representation: 70 Years of Bollywood
- Independent Ethical Assessment of Text Classification Models: A Hate Speech Detection Case Study
- Assessing the Reliability of Word Embedding Gender Bias Measures
- Low-skilled Occupations Face the Highest Upskilling Pressure
- A Source-Criticism Debiasing Method for GloVe Embeddings
- Gender Bias Hidden Behind Chinese Word Embeddings: The Case of Chinese Adjectives
- The presence of occupational structure in online texts based on word embedding NLP models
- From communities to interpretable network and word embedding: an unified approach
- Interpretable Word Embeddings via Informative Priors
- Automatically Inferring Gender Associations from Language
- Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs
- The Golden Rule as a Heuristic to Measure the Fairness of Texts Using Machine Learning
- A Computational Social Science Approach to Understanding Predictors of Chafee Service Receipt
- From Symbols to Embeddings: A Tale of Two Representations in Computational Social Science
- Parallax: Visualizing and Understanding the Semantics of Embedding Spaces via Algebraic Formulae
- A Probabilistic Framework for Learning Domain Specific Hierarchical Word Embeddings
- UnibucKernel: Geolocating Swiss German Jodels Using Ensemble Learning
- How Does Fine-tuning Affect the Geometry of Embedding Space: A Case Study on Isotropy
- Semantic coordinates analysis reveals language changes in the AI field
- Unpacking the Interdependent Systems of Discrimination: Ableist Bias in NLP Systems through an Intersectional Lens
- Using Sociolinguistic Variables to Reveal Changing Attitudes Towards Sexuality and Gender
- Sexism in the Judiciary
- Detecting Cross-Geographic Biases in Toxicity Modeling on Social Media
- Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors
- Gender Stereotype Reinforcement: Measuring the Gender Bias Conveyed by Ranking Algorithms
- LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data
- Measuring a Texts Fairness Dimensions Using Machine Learning Based on Social Psychological Factors
- Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji Embeddings
- An Improved Historical Embedding without Alignment