Considerations When Learning Additive Explanations for Black-Box Models
arXiv:1801.08640 · doi:10.1007/s10994-023-06335-8
Abstract
Many methods to explain black-box models, whether local or global, are additive. In this paper, we study global additive explanations for non-additive models, focusing on four explanation methods: partial dependence, Shapley explanations adapted to a global setting, distilled additive explanations, and gradient-based explanations. We show that different explanation methods characterize non-additive components in a black-box model's prediction function in different ways. We use the concepts of main and total effects to anchor additive explanations, and quantitatively evaluate additive and non-additive explanations. Even though distilled explanations are generally the most accurate additive explanations, non-additive explanations such as tree explanations that explicitly model non-additive components tend to be even more accurate. Despite this, our user study showed that machine learning practitioners were better able to leverage additive explanations for various tasks. These considerations should be taken into account when considering which explanation to trust and use to explain black-box models.
Published at Machine Learning (2023). Previously titled "Learning Global Additive Explanations for Neural Nets Using Model Distillation". A short version was presented at NeurIPS 2018 Machine Learning for Health Workshop
References in corpus (19)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Towards A Rigorous Science of Interpretable Machine Learning
- Methods for Interpreting and Understanding Deep Neural Networks
- Predictive learning via rule ensembles
- All Models are Wrong, but Many are Useful: Learning a Variable's Importance by Studying an Entire Class of Prediction Models Simultaneously
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- InterpretML: A Unified Framework for Machine Learning Interpretability
- Distilling a Neural Network Into a Soft Decision Tree
- Prototype selection for interpretable classification
- Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods
- GLocalX -- From Local to Global Explanations of Black Box AI Models
- Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation
- How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations
- Global Aggregations of Local Explanations for Black Box models
- Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses
- Global Explanations of Neural Networks: Mapping the Landscape of Predictions
- Efficient nonparametric statistical inference on population feature importance using Shapley values