Towards Auditability for Fairness in Deep Learning
arXiv:2012.00106
Abstract
Group fairness metrics can detect when a deep learning model behaves differently for advantaged and disadvantaged groups, but even models that score well on these metrics can make blatantly unfair predictions. We present smooth prediction sensitivity, an efficiently computed measure of individual fairness for deep learning models that is inspired by ideas from interpretability in deep learning. smooth prediction sensitivity allows individual predictions to be audited for fairness. We present preliminary experimental results suggesting that smooth prediction sensitivity can help distinguish between fair and unfair predictions, and that it may be helpful in detecting blatantly unfair predictions from "group-fair" models.
Presented at the workshop on Algorithmic Fairness through the Lens of Causality and Interpretability (AFCI'20)
References in corpus (7)
- Equality of Opportunity in Supervised Learning
- SmoothGrad: removing noise by adding noise
- Data Decisions and Theoretical Implications when Adversarially Learning Fair Representations
- No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
- A statistical framework for fair predictive algorithms
- Improved Adversarial Learning for Fair Classification
- Towards a Measure of Individual Fairness for Deep Learning