A survey on measuring indirect discrimination in machine learning
arXiv:1511.00148
Abstract
Nowadays, many decisions are made using predictive models built on historical data.Predictive models may systematically discriminate groups of people even if the computing process is fair and well-intentioned. Discrimination-aware data mining studies how to make predictive models free from discrimination, when historical data, on which they are built, may be biased, incomplete, or even contain past discriminatory decisions. Discrimination refers to disadvantageous treatment of a person based on belonging to a category rather than on individual merit. In this survey we review and organize various discrimination measures that have been used for measuring discrimination in data, as well as in evaluating performance of discrimination-aware predictive models. We also discuss related measures from other disciplines, which have not been used for measuring discrimination, but potentially could be suitable for this purpose. We computationally analyze properties of selected measures. We also review and discuss measuring procedures, and present recommendations for practitioners. The primary target audience is data mining, machine learning, pattern recognition, statistical modeling researchers developing new methods for non-discriminatory predictive modeling. In addition, practitioners and policy makers would use the survey for diagnosing potential discrimination by predictive models.
References in corpus (1)
Cited by in corpus (36)
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
- On the (im)possibility of fairness
- On Formalizing Fairness in Prediction with Machine Learning
- FairSight: Visual Analytics for Fairness in Decision Making
- Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints
- Fairness in Algorithmic Decision Making: An Excursion Through the Lens of Causality
- Predicting Demographics, Moral Foundations, and Human Values from Digital Behaviors
- Quantifying and Reducing Stereotypes in Word Embeddings
- Machine learning fairness notions: Bridging the gap with real-world applications
- Synthetic Data for Social Good
- Provably Fair Representations
- Visus: An Interactive System for Automatic Machine Learning Model Building and Curation
- Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems
- Fairness-Aware Learning with Prejudice Free Representations
- FAHT: An Adaptive Fairness-aware Decision Tree Classifier
- Auditing and Achieving Intersectional Fairness in Classification Problems
- Research Directions for Principles of Data Management (Dagstuhl Perspectives Workshop 16151)
- Evaluating Fairness Metrics in the Presence of Dataset Bias
- Online and Customizable Fairness-aware Learning
- Assessing Fairness in Classification Parity of Machine Learning Models in Healthcare
- Detection and Mitigation of Bias in Ted Talk Ratings
- Is there Gender bias and stereotype in Portuguese Word Embeddings?
- Fairness Warnings and Fair-MAML: Learning Fairly with Minimal Data
- A Normative approach to Attest Digital Discrimination
- Responsible and Representative Multimodal Data Acquisition and Analysis: On Auditability, Benchmarking, Confidence, Data-Reliance & Explainability
- A Primal-Dual Subgradient Approachfor Fair Meta Learning
- Pooling of Causal Models under Counterfactual Fairness via Causal Judgement Aggregation
- Integral Privacy for Sampling
- TFW, DamnGina, Juvie, and Hotsie-Totsie: On the Linguistic and Social Aspects of Internet Slang
- FARF: A Fair and Adaptive Random Forests Classifier
- Fair Meta-Learning For Few-Shot Classification
- Unfairness Discovery and Prevention For Few-Shot Regression
- Rank-Based Multi-task Learning for Fair Regression
- Data Management for Causal Algorithmic Fairness
- Counterfactually Fair Prediction Using Multiple Causal Models
- Is it ethical to avoid error analysis?