Measuring and improving the quality of visual explanations
arXiv:2003.08774
Abstract
The ability of to explain neural network decisions goes hand in hand with their safe deployment. Several methods have been proposed to highlight features important for a given network decision. However, there is no consensus on how to measure effectiveness of these methods. We propose a new procedure for evaluating explanations. We use it to investigate visual explanations extracted from a range of possible sources in a neural network. We quantify the benefit of combining these sources and challenge a recent appeal for taking bias parameters into account. We support our conclusions with a general assessment of the impact of bias parameters in ImageNet classifiers
References in corpus (9)
- Distilling the Knowledge in a Neural Network
- Axiomatic Attribution for Deep Networks
- Learning Important Features Through Propagating Activation Differences
- SmoothGrad: removing noise by adding noise
- Sanity Checks for Saliency Maps
- A Benchmark for Interpretability Methods in Deep Neural Networks
- Benchmarking Attribution Methods with Relative Feature Importance
- Full-Gradient Representation for Neural Network Visualization
- Aggregating explanation methods for stable and robust explainability