On Interpretability of Artificial Neural Networks: A Survey
arXiv:2001.02522
Abstract
Deep learning as represented by the artificial deep neural networks (DNNs) has achieved great success in many important areas that deal with text, images, videos, graphs, and so on. However, the black-box nature of DNNs has become one of the primary obstacles for their wide acceptance in mission-critical applications such as medical diagnosis and therapy. Due to the huge potential of deep learning, interpreting neural networks has recently attracted much research attention. In this paper, based on our comprehensive taxonomy, we systematically review recent studies in understanding the mechanism of neural networks, describe applications of interpretability especially in medicine, and discuss future directions of interpretability research, such as in relation to fuzzy logic and brain science.
References in corpus (47)
- Distilling the Knowledge in a Neural Network
- Semi-Supervised Classification with Graph Convolutional Networks
- Attention U-Net: Learning Where to Look for the Pancreas
- Towards A Rigorous Science of Interpretable Machine Learning
- Striving for Simplicity: The All Convolutional Net
- Confounding variables can degrade generalization performance of radiological deep learning models
- Reconciling modern machine learning practice and the bias-variance trade-off
- Understanding Neural Networks Through Deep Visualization
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- SmoothGrad: removing noise by adding noise
- Object Detectors Emerge in Deep Scene CNNs
- Sanity Checks for Saliency Maps
- This Looks Like That: Deep Learning for Interpretable Image Recognition
- Neural Ordinary Differential Equations
- Attention is not Explanation
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- Understanding Neural Networks through Representation Erasure
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- Can Deep Learning Outperform Modern Commercial CT Image Reconstruction Methods?
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Understanding the Role of Individual Units in a Deep Neural Network
- Prototype selection for interpretable classification
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Local Rule-Based Explanations of Black Box Decision Systems
- Unsupervised Learning by Competing Hidden Units
- Why Deep Neural Networks for Function Approximation?
- An exact mapping between the Variational Renormalization Group and Deep Learning
- ResNet with one-neuron hidden layers is a Universal Approximator
- Beyond Sparsity: Tree Regularization of Deep Models for Interpretability
- ISeeU: Visually interpretable deep learning for mortality prediction inside the ICU
- Hierarchical interpretations for neural network predictions
- Shape and Margin-Aware Lung Nodule Classification in Low-dose CT Images via Soft Activation Mapping
- Explainable Neural Networks based on Additive Index Models
- Norm-Based Capacity Control in Neural Networks
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- An Interpretable Model with Globally Consistent Explanations for Credit Risk
- Restricting the Flow: Information Bottlenecks for Attribution
- Adapted Deep Embeddings: A Synthesis of Methods for -Shot Inductive Transfer Learning
- Identifying Unknown Unknowns in the Open World: Representations and Policies for Guided Exploration
- Counterfactual Visual Explanations
- Interpreting Adversarial Examples by Activation Promotion and Suppression
- Mechanisms of dimensionality reduction and decorrelation in deep neural networks
- Producing radiologist-quality reports for interpretable artificial intelligence
- Graph Structure of Neural Networks
- Understanding Impacts of High-Order Loss Approximations and Features in Deep Learning Interpretation
- Towards Interpretable R-CNN by Unfolding Latent Structures
- Neural Networks, Hypersurfaces, and Radon Transforms
Cited by in corpus (4)
- Advancing from Predictive Maintenance to Intelligent Maintenance with AI and IIoT
- VoxelHop: Successive Subspace Learning for ALS Disease Classification Using Structural MRI
- Ada-SISE: Adaptive Semantic Input Sampling for Efficient Explanation of Convolutional Neural Networks
- Understanding in Artificial Intelligence