When Explanations Lie: Why Many Modified BP Attributions Fail
arXiv:1912.09818
Abstract
Attribution methods aim to explain a neural network's prediction by highlighting the most relevant image areas. A popular approach is to backpropagate (BP) a custom relevance score using modified rules, rather than the gradient. We analyze an extensive set of modified BP methods: Deep Taylor Decomposition, Layer-wise Relevance Propagation (LRP), Excitation BP, PatternAttribution, DeepLIFT, Deconv, RectGrad, and Guided BP. We find empirically that the explanations of all mentioned methods, except for DeepLIFT, are independent of the parameters of later layers. We provide theoretical insights for this surprising behavior and also analyze why DeepLIFT does not suffer from this limitation. Empirically, we measure how information of later layers is ignored by using our new metric, cosine similarity convergence (CSC). The paper provides a framework to assess the faithfulness of new and existing modified BP methods theoretically and empirically. For code see: https://github.com/berleon/when-explanations-lie
Published in ICML 2020 - Updated Proof
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Towards A Rigorous Science of Interpretable Machine Learning
- SmoothGrad: removing noise by adding noise
- Investigating the influence of noise and distractors on the interpretation of neural networks
- Restricting the Flow: Information Bottlenecks for Attribution
Cited by in corpus (5)
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
- Harnessing spatial homogeneity of neuroimaging data: patch individual filter layers for CNNs
- Saliency strikes back: How filtering out high frequencies improves white-box explanations
- Visual Probing: Cognitive Framework for Explaining Self-Supervised Image Representations