Learning how to explain neural networks: PatternNet and PatternAttribution
arXiv:1705.05598
Abstract
DeConvNet, Guided BackProp, LRP, were invented to better understand deep neural networks. We show that these methods do not produce the theoretically correct explanation for a linear model. Yet they are used on multi-layer networks with millions of parameters. This is a cause for concern since linear models are simple neural networks. We argue that explanation methods for neural nets should work reliably in the limit of simplicity, the linear models. Based on our analysis of linear models we propose a generalization that yields two explanation techniques (PatternNet and PatternAttribution) that are theoretically sound for linear models and produce improved explanations for deep networks.
Cited by in corpus (81)
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- SchNet - a deep learning architecture for molecules and materials
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Towards Explainable Artificial Intelligence
- An Explainable 3D Residual Self-Attention Deep Neural Network FOR Joint Atrophy Localization and Alzheimer's Disease Diagnosis using Structural MRI
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- Explanations can be manipulated and geometry is to blame
- NeuralHydrology -- Interpreting LSTMs in Hydrology
- An Evaluation of the Human-Interpretability of Explanation
- Neural Network Attribution Methods for Problems in Geoscience: A Novel Synthetic Benchmark Dataset
- How do Humans Understand Explanations from Machine Learning Systems? An Evaluation of the Human-Interpretability of Explanation
- Explainable Artificial Intelligence: a Systematic Review
- Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Explainability Techniques for Graph Convolutional Networks
- A Survey of Deep Learning for Scientific Discovery
- Testing and verification of neural-network-based safety-critical control software: A systematic literature review
- Restricting the Flow: Information Bottlenecks for Attribution
- A Theoretical Explanation for Perplexing Behaviors of Backpropagation-based Visualizations
- Beyond saliency: understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation
- Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Values Approximation
- Survey of XAI in digital pathology
- When Explanations Lie: Why Many Modified BP Attributions Fail
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- iNNvestigate neural networks!
- Interpreting CNNs via Decision Trees
- Interpretable deep learning for nuclear deformation in heavy ion collisions
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Explaining Convolutional Neural Networks using Softmax Gradient Layer-wise Relevance Propagation
- Global Explanations of Neural Networks: Mapping the Landscape of Predictions
- Knowledge Consistency between Neural Networks and Beyond
- Explanation-Guided Training for Cross-Domain Few-Shot Classification
- Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
- Explaining Knowledge Distillation by Quantifying the Knowledge
- Deep Learning Development Environment in Virtual Reality
- Deeply Explain CNN via Hierarchical Decomposition
- Feature Perturbation Augmentation for Reliable Evaluation of Importance Estimators in Neural Networks
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural Networks
- Explaining Bayesian Neural Networks
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- AUTOLYCUS: Exploiting Explainable AI (XAI) for Model Extraction Attacks against Interpretable Models
- SoK: Machine Learning Governance
- Trustworthy Convolutional Neural Networks: A Gradient Penalized-based Approach
- A Baseline for Shapley Values in MLPs: from Missingness to Neutrality
- Debugging Tests for Model Explanations
- A Game-Theoretic Taxonomy of Visual Concepts in DNNs
- Towards a Unified Evaluation of Explanation Methods without Ground Truth
- XProtoNet: Diagnosis in Chest Radiography with Global and Local Explanations
- XRAI: Better Attributions Through Regions
- Human-Expert-Level Brain Tumor Detection Using Deep Learning with Data Distillation and Augmentation
- Towards Robust Explanations for Deep Neural Networks
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Shapley Interpretation and Activation in Neural Networks
- Explaining AlphaGo: Interpreting Contextual Effects in Neural Networks
- Data-Adaptive Discriminative Feature Localization with Statistically Guaranteed Interpretation
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- Progress in deep Markov State Modeling: Coarse graining and experimental data restraints
- Discriminative Attribution from Counterfactuals
- Improved Feature Importance Computations for Tree Models: Shapley vs. Banzhaf
- Understanding of Kernels in CNN Models by Suppressing Irrelevant Visual Features in Images
- Simplifying the explanation of deep neural networks with sufficient and necessary feature-sets: case of text classification
- Unifying Model Explainability and Robustness via Machine-Checkable Concepts
- The Partial Response Network: a neural network nomogram
- Weakly-Supervised Cell Tracking via Backward-and-Forward Propagation
- Pattern-Guided Integrated Gradients
- Understanding Regularization to Visualize Convolutional Neural Networks
- Understanding Recurrent Neural State Using Memory Signatures
- Learning Shape Features and Abstractions in 3D Convolutional Neural Networks for Detecting Alzheimer's Disease
- Analysis of Atomistic Representations Using Weighted Skip-Connections
- An Overview of Computational Approaches for Interpretation Analysis
- Explain and Improve: LRP-Inference Fine-Tuning for Image Captioning Models
- Attribution Mask: Filtering Out Irrelevant Features By Recursively Focusing Attention on Inputs of DNNs
- Learning Propagation Rules for Attribution Map Generation
- Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop
- Explainable Deep Modeling of Tabular Data using TableGraphNet
- Verifiability and Predictability: Interpreting Utilities of Network Architectures for Point Cloud Processing
- Staging Epileptogenesis with Deep Neural Networks
- A copula-based visualization technique for a neural network