On the Robustness of Interpretability Methods
arXiv:1806.08049
Abstract
We argue that robustness of explanations---i.e., that similar inputs should give rise to similar explanations---is a key desideratum for interpretability. We introduce metrics to quantify robustness and demonstrate that current methods do not perform well according to these metrics. Finally, we propose ways that robustness can be enforced on existing interpretability approaches.
presented at 2018 ICML Workshop on Human Interpretability in Machine Learning (WHI 2018), Stockholm, Sweden
References in corpus (3)
Cited by in corpus (21)
- Statistical stability indices for LIME: obtaining reliable explanations for Machine Learning models
- Robust Explainability: A Tutorial on Gradient-Based Attribution Methods for Deep Neural Networks
- Acquisition of Chess Knowledge in AlphaZero
- Counterfactual Explanations for Machine Learning on Multivariate Time Series Data
- Explainability in Music Recommender Systems
- A survey of algorithmic recourse: definitions, formulations, solutions, and prospects
- Regularizing Black-box Models for Improved Interpretability
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- BayLIME: Bayesian Local Interpretable Model-Agnostic Explanations
- Harnessing the Vulnerability of Latent Layers in Adversarially Trained Models
- Interpretable Machine Learning: Moving From Mythos to Diagnostics
- This changes to that : Combining causal and non-causal explanations to generate disease progression in capsule endoscopy
- MeLIME: Meaningful Local Explanation for Machine Learning Models
- An exact counterfactual-example-based approach to tree-ensemble models interpretability
- Explaining Predictions from Tree-based Boosting Ensembles
- Better sampling in explanation methods can prevent dieselgate-like deception
- Regularizing Black-box Models for Improved Interpretability (HILL 2019 Version)
- Logic Traps in Evaluating Attribution Scores
- Scaling Symbolic Methods using Gradients for Neural Model Explanation
- XPROAX-Local explanations for text classification with progressive neighborhood approximation
- Defense Against Explanation Manipulation