Identifying and Correcting Label Bias in Machine Learning
arXiv:1901.04966
Abstract
Datasets often contain biases which unfairly disadvantage certain groups, and classifiers trained on such datasets can inherit these biases. In this paper, we provide a mathematical formulation of how this bias can arise. We do so by assuming the existence of underlying, unknown, and unbiased labels which are overwritten by an agent who intends to provide accurate labels but may have biases against certain groups. Despite the fact that we only observe the biased labels, we are able to show that the bias may nevertheless be corrected by re-weighting the data points without changing the labels. We show, with theoretical guarantees, that training on the re-weighted dataset corresponds to training on the unobserved but unbiased labels, thus leading to an unbiased machine learning classifier. Our procedure is fast and robust and can be used with virtually any learning algorithm. We evaluate on a number of standard machine learning fairness datasets and a variety of fairness notions, finding that our method outperforms standard approaches in achieving fair classification.
References in corpus (4)
Cited by in corpus (13)
- Does Object Recognition Work for Everyone?
- Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation
- On Adversarial Bias and the Robustness of Fair Machine Learning
- Counterfactual Reasoning for Fair Clinical Risk Prediction
- Benign Shortcut for Debiasing: Fair Visual Recognition via Intervention with Shortcut Features
- Beyond Low Earth Orbit: Biomonitoring, Artificial Intelligence, and Precision Space Health
- Controlling biases and diversity in diverse image-to-image translation
- Physics Enhanced Artificial Intelligence
- Robust Semantic Interpretability: Revisiting Concept Activation Vectors
- An Empirical Investigation of Learning from Biased Toxicity Labels
- Doing good by fighting fraud: Ethical anti-fraud systems for mobile payments
- Small Business Classification By Name: Addressing Gender and Geographic Origin Biases
- Group-based Fair Learning Leads to Counter-intuitive Predictions