Embedding Deep Networks into Visual Explanations
arXiv:1709.05360 · doi:10.1016/j.artint.2020.103435
Abstract
In this paper, we propose a novel Explanation Neural Network (XNN) to explain the predictions made by a deep network. The XNN works by learning a nonlinear embedding of a high-dimensional activation vector of a deep network layer into a low-dimensional explanation space while retaining faithfulness i.e., the original deep learning predictions can be constructed from the few concepts extracted by our explanation network. We then visualize such concepts for human to learn about the high-level concepts that the deep network is using to make decisions. We propose an algorithm called Sparse Reconstruction Autoencoder (SRAE) for learning the embedding to the explanation space. SRAE aims to reconstruct part of the original feature space while retaining faithfulness. A pull-away term is applied to SRAE to make the bases of the explanation space more orthogonal to each other. A visualization system is then introduced for human understanding of the features in the explanation space. The proposed method is applied to explain CNN models in image classification tasks. We conducted a human study, which shows that the proposed approach outperforms single saliency map baselines, and improves human performance on a difficult classification tasks. Also, several novel metrics are introduced to evaluate the performance of explanations quantitatively without human involvement.
References in corpus (11)
- Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
- Towards A Rigorous Science of Interpretable Machine Learning
- Energy-based Generative Adversarial Network
- Compressing Neural Networks with the Hashing Trick
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Measuring the Quality of Explanations: The System Causability Scale (SCS). Comparing Human and Machine Explanations
- RISE: Randomized Input Sampling for Explanation of Black-box Models
- The (Un)reliability of saliency methods
- Deep Learning for Case-Based Reasoning through Prototypes: A Neural Network that Explains Its Predictions
- Visual Explanation by Interpretation: Improving Visual Feedback Capabilities of Deep Neural Networks
- The morphodynamics of 3D migrating cancer cells
Cited by in corpus (5)
- From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI
- Causal Learning and Explanation of Deep Neural Networks via Autoencoded Activations
- Counterfactual Visual Explanations
- How explainable AI affects human performance: A systematic review of the behavioural consequences of saliency maps
- From Heatmaps to Structural Explanations of Image Classifiers