Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
arXiv:1901.09451 · doi:10.1145/3287560.3287572
Abstract
We present a large-scale study of gender bias in occupation classification, a task where the use of machine learning may lead to negative outcomes on peoples' lives. We analyze the potential allocation harms that can result from semantic representation bias. To do so, we study the impact on occupation classification of including explicit gender indicators---such as first names and pronouns---in different semantic representations of online biographies. Additionally, we quantify the bias that remains when these indicators are "scrubbed," and describe proxy behavior that occurs in the absence of explicit gender indicators. As we demonstrate, differences in true positive rates between genders are correlated with existing gender imbalances in occupations, which may compound these imbalances.
Accepted at ACM Conference on Fairness, Accountability, and Transparency (ACM FAT*), 2019
Cited by in corpus (29)
- Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale
- Representation Bias in Data: A Survey on Identification and Resolution Techniques
- Fairness and Bias in Algorithmic Hiring: a Multidisciplinary Survey
- A Meta-Analysis of the Utility of Explainable Artificial Intelligence in Human-AI Decision-Making
- Gender Bias in Word Embeddings: A Comprehensive Analysis of Frequency, Syntax, and Semantics
- Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making
- Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment
- Societal Biases in Retrieved Contents: Measurement Framework and Adversarial Mitigation for BERT Rankers
- A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
- Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
- Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time
- Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
- Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
- Justice in Misinformation Detection Systems: An Analysis of Algorithms, Stakeholders, and Potential Harms
- LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions
- Measuring justice in machine learning
- Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- An investigation of structures responsible for gender bias in BERT and DistilBERT
- Social Norm Bias: Residual Harms of Fairness-Aware Algorithms
- Unmasking Gender Bias in Recommendation Systems and Enhancing Category-Aware Fairness
- Don't let Ricci v. DeStefano Hold You Back: A Bias-Aware Legal Solution to the Hiring Paradox
- NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
- Bias-Aware Agent: Enhancing Fairness in AI-Driven Knowledge Retrieval
- Fairness Definitions in Language Models Explained
- Algorithmic Fairness Datasets: the Story so Far
- Interpretable by Design: Learning Predictors by Composing Interpretable Queries
- A Novel Information-Theoretic Objective to Disentangle Representations for Fair Classification
- Conceptualizing Uncertainty: A Concept-based Approach to Explaining Uncertainty