The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations
arXiv:2205.03295 · doi:10.1145/3531146.3533179
Abstract
Machine learning models in safety-critical settings like healthcare are often blackboxes: they contain a large number of parameters which are not transparent to users. Post-hoc explainability methods where a simple, human-interpretable model imitates the behavior of these blackbox models are often proposed to help users trust model predictions. In this work, we audit the quality of such explanations for different protected subgroups using real data from four settings in finance, healthcare, college admissions, and the US justice system. Across two different blackbox model architectures and four popular explainability methods, we find that the approximation quality of explanation models, also known as the fidelity, differs significantly between subgroups. We also demonstrate that pairing explainability methods with recent advances in robust machine learning can improve explanation fairness in some settings. However, we highlight the importance of communicating details of non-zero fidelity gaps to users, since a single solution might not exist across all settings. Finally, we discuss the implications of unfair explanation models as a challenging and understudied problem facing the machine learning community.
Published in FAccT 2022
References in corpus (7)
- A Survey on the Explainability of Supervised Machine Learning
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems
- Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning
- Automated Machine Learning in Practice: State of the Art and Recent Results
- Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
- Pulling Up by the Causal Bootstraps: Causal Data Augmentation for Pre-training Debiasing
Cited by in corpus (4)
- Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
- Evaluating the Impact of Social Determinants on Health Prediction in the Intensive Care Unit
- In defence of post-hoc explanations in medical AI
- GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations