Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation
arXiv:1710.06169 · doi:10.1145/3278721.3278725
Abstract
Black-box risk scoring models permeate our lives, yet are typically proprietary or opaque. We propose Distill-and-Compare, a model distillation and comparison approach to audit such models. To gain insight into black-box models, we treat them as teachers, training transparent student models to mimic the risk scores assigned by black-box models. We compare the student model trained with distillation to a second un-distilled transparent model trained on ground-truth outcomes, and use differences between the two models to gain insight into the black-box model. Our approach can be applied in a realistic setting, without probing the black-box model API. We demonstrate the approach on four public data sets: COMPAS, Stop-and-Frisk, Chicago Police, and Lending Club. We also propose a statistical test to determine if a data set is missing key features used to train the black-box model. Our test finds that the ProPublica data is likely missing key feature(s) used in COMPAS.
Camera-ready version for AAAI/ACM AIES 2018. Data and pseudocode at https://github.com/shftan/auditblackbox. Previously titled "Detecting Bias in Black-Box Models Using Transparent Model Distillation". A short version was presented at NIPS 2017 Symposium on Interpretable Machine Learning
References in corpus (5)
Cited by in corpus (12)
- The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies
- How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations
- Denoising diffusion algorithm for inverse design of microstructures with fine-tuned nonlinear material properties
- The Who in XAI: How AI Background Shapes Perceptions of AI Explanations
- Explainable Diabetic Retinopathy Detection and Retinal Image Generation
- The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations
- Considerations When Learning Additive Explanations for Black-Box Models
- The Bouncer Problem: Challenges to Remote Explainability
- Why Should I Choose You? AutoXAI: A Framework for Selecting and Tuning eXplainable AI Solutions
- Reflective-Net: Learning from Explanations
- Accurate Explanation Model for Image Classifiers using Class Association Embedding
- Access Denied: Meaningful Data Access for Quantitative Algorithm Audits