The (Un)reliability of saliency methods
arXiv:1711.00867
Abstract
Saliency methods aim to explain the predictions of deep neural networks. These methods lack reliability when the explanation is sensitive to factors that do not contribute to the model prediction. We use a simple and common pre-processing step ---adding a constant shift to the input data--- to show that a transformation with no effect on the model can cause numerous methods to incorrectly attribute. In order to guarantee reliability, we posit that methods should fulfill input invariance, the requirement that a saliency method mirror the sensitivity of the model with respect to transformations of the input. We show, through several examples, that saliency methods that do not satisfy input invariance result in misleading attribution.
Cited by in corpus (12)
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Neural Network Attributions: A Causal Perspective
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- Interpretable deep learning for nuclear deformation in heavy ion collisions
- Global Explanations of Neural Networks: Mapping the Landscape of Predictions
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- Semantics for Global and Local Interpretation of Deep Neural Networks
- What Do Adversarially Robust Models Look At?
- Regional Image Perturbation Reduces Norms of Adversarial Examples While Maintaining Model-to-model Transferability
- Regression Concept Vectors for Bidirectional Explanations in Histopathology
- Games for Fairness and Interpretability