Saliency Methods for Explaining Adversarial Attacks
arXiv:1908.08413
Abstract
The classification decisions of neural networks can be misled by small imperceptible perturbations. This work aims to explain the misled classifications using saliency methods. The idea behind saliency methods is to explain the classification decisions of neural networks by creating so-called saliency maps. Unfortunately, a number of recent publications have shown that many of the proposed saliency methods do not provide insightful explanations. A prominent example is Guided Backpropagation (GuidedBP), which simply performs (partial) image recovery. However, our numerical analysis shows the saliency maps created by GuidedBP do indeed contain class-discriminative information. We propose a simple and efficient way to enhance the saliency maps. The proposed enhanced GuidedBP shows the state-of-the-art performance to explain adversary classifications.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Learning Important Features Through Propagating Activation Differences
- SmoothGrad: removing noise by adding noise
- Sanity Checks for Saliency Maps
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- Explanations can be manipulated and geometry is to blame
Cited by in corpus (9)
- Attacking Adversarial Attacks as A Defense
- Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
- Interpretable Graph Capsule Networks for Object Recognition
- Contextual Prediction Difference Analysis for Explaining Individual Image Classifications
- Visualizing Automatic Speech Recognition -- Means for a Better Understanding?
- Identifying Untrustworthy Predictions in Neural Networks by Geometric Gradient Analysis
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Heat and Blur: An Effective and Fast Defense Against Adversarial Examples
- Understanding Robustness in Teacher-Student Setting: A New Perspective