Explaining NonLinear Classification Decisions with Deep Taylor Decomposition
arXiv:1512.02479 · doi:10.1016/j.patcog.2016.11.008
Abstract
Nonlinear methods such as Deep Neural Networks (DNNs) are the gold standard for various challenging machine learning problems, e.g., image classification, natural language processing or human action recognition. Although these methods perform impressively well, they have a significant disadvantage, the lack of transparency, limiting the interpretability of the solution and thus the scope of application in practice. Especially DNNs act as black boxes due to their multilayer nonlinear structure. In this paper we introduce a novel methodology for interpreting generic multilayer neural networks by decomposing the network classification decision into contributions of its input elements. Although our focus is on image classification, the method is applicable to a broad set of input data, learning tasks and network architectures. Our method is based on deep Taylor decomposition and efficiently utilizes the structure of the network by backpropagating the explanations from the output to the input layer. We evaluate the proposed method empirically on the MNIST and ILSVRC data sets.
20 pages, 15 figures
References in corpus (4)
Cited by in corpus (268)
- A Survey on Deep Learning in Medical Image Analysis
- Deep learning with convolutional neural networks for EEG decoding and visualization
- Methods for Interpreting and Understanding Deep Neural Networks
- SchNet - a deep learning architecture for molecules and materials
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
- A Survey on the Explainability of Supervised Machine Learning
- Combining Machine Learning and Computational Chemistry for Predictive Insights Into Chemical Systems
- The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies
- 'It's Reducing a Human Being to a Percentage'; Perceptions of Justice in Algorithmic Decisions
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Towards Explainable Artificial Intelligence
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- Fairness and Accountability Design Needs for Algorithmic Support in High-Stakes Public Sector Decision-Making
- Analyzing Federated Learning through an Adversarial Lens
- Insightful classification of crystal structures using deep learning
- Physically Interpretable Neural Networks for the Geosciences: Applications to Earth System Variability
- SNAS: Stochastic Neural Architecture Search
- "What is Relevant in a Text Document?": An Interpretable Machine Learning Approach
- Higher-Order Explanations of Graph Neural Networks via Relevant Walks
- Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
- Explaining the Unique Nature of Individual Gait Patterns with Deep Learning
- explAIner: A Visual Analytics Framework for Interactive and Explainable Machine Learning
- Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future Direction
- An Explainable 3D Residual Self-Attention Deep Neural Network FOR Joint Atrophy Localization and Alzheimer's Disease Diagnosis using Structural MRI
- A Review on Explainable Artificial Intelligence for Healthcare: Why, How, and When?
- A Deep Learning Interpretable Classifier for Diabetic Retinopathy Disease Grading
- Ablation Studies in Artificial Neural Networks
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- Resolving challenges in deep learning-based analyses of histopathological images using explanation methods
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Explanations can be manipulated and geometry is to blame
- Enslaving the Algorithm: From a "Right to an Explanation" to a "Right to Better Decisions"?
- NeuralHydrology -- Interpreting LSTMs in Hydrology
- Evaluation of post-hoc interpretability methods in time-series classification
- Post-hoc explanation of black-box classifiers using confident itemsets
- Explainable artificial intelligence in breast cancer detection and risk prediction: A systematic scoping review
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications
- From Clustering to Cluster Explanations via Neural Networks
- Robust Explainability: A Tutorial on Gradient-Based Attribution Methods for Deep Neural Networks
- Explainable Deep One-Class Classification
- Neural Network Attribution Methods for Problems in Geoscience: A Novel Synthetic Benchmark Dataset
- RUDDER: Return Decomposition for Delayed Rewards
- Towards Explaining Anomalies: A Deep Taylor Decomposition of One-Class Models
- Elucidating the Behavior of Nanophotonic Structures Through Explainable Machine Learning Algorithms
- Investigating the influence of noise and distractors on the interpretation of neural networks
- Explainable Artificial Intelligence: a Systematic Review
- Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience
- Deep Divergence-Based Approach to Clustering
- Explainable AI: current status and future directions
- On the Explanation of Machine Learning Predictions in Clinical Gait Analysis
- Explaining and Interpreting LSTMs
- Interpretation of Neural Networks is Fragile
- Indicator patterns of forced change learned by an artificial neural network
- Probing slow earthquakes with deep learning
- Building and Interpreting Deep Similarity Models
- Neuron Shapley: Discovering the Responsible Neurons
- Towards Relatable Explainable AI with the Perceptual Process
- Explainability Techniques for Graph Convolutional Networks
- Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization
- Testing and verification of neural-network-based safety-critical control software: A systematic literature review
- Restricting the Flow: Information Bottlenecks for Attribution
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Beyond saliency: understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation
- Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Values Approximation
- Towards Evaluating Explanations of Vision Transformers for Medical Imaging
- Analysis of a Deep Learning Model for 12-Lead ECG Classification Reveals Learned Features Similar to Diagnostic Criteria
- On Interpretability of Artificial Neural Networks: A Survey
- Stakeholders in Explainable AI
- When Explanations Lie: Why Many Modified BP Attributions Fail
- Survey of XAI in digital pathology
- DeepCOVIDExplainer: Explainable COVID-19 Diagnosis Based on Chest X-ray Images
- Saccader: Improving Accuracy of Hard Attention Models for Vision
- NormLime: A New Feature Importance Metric for Explaining Deep Neural Networks
- Iterative Augmentation of Visual Evidence for Weakly-Supervised Lesion Localization in Deep Interpretability Frameworks: Application to Color Fundus Images
- Supporting DNN Safety Analysis and Retraining through Heatmap-based Unsupervised Learning
- Understanding and Comparing Deep Neural Networks for Age and Gender Classification
- Robust and interpretable blind image denoising via bias-free convolutional neural networks
- Gradual Channel Pruning while Training using Feature Relevance Scores for Convolutional Neural Networks
- Neural Network Attributions: A Causal Perspective
- Granger-causal Attentive Mixtures of Experts: Learning Important Features with Neural Networks
- Understanding Integrated Gradients with SmoothTaylor for Deep Neural Network Attribution
- White Box Methods for Explanations of Convolutional Neural Networks in Image Classification Tasks
- Neural Machine Reading Comprehension: Methods and Trends
- RES: A Robust Framework for Guiding Visual Explanation
- Counterfactual Explanation Based on Gradual Construction for Deep Networks
- Explaining Clinical Decision Support Systems in Medical Imaging using Cycle-Consistent Activation Maximization
- Multi-objective optimization determines when, which and how to fuse deep networks: an application to predict COVID-19 outcomes
- Vulnerabilities of Connectionist AI Applications: Evaluation and Defence
- Explaining Black-box Models for Biomedical Text Classification
- BrainNPT: Pre-training of Transformer networks for brain network classification
- Interpreting the Predictions of Complex ML Models by Layer-wise Relevance Propagation
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Interpretable deep learning for nuclear deformation in heavy ion collisions
- Self-Explaining Structures Improve NLP Models
- Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers
- An explainable deep vision system for animal classification and detection in trail-camera images with automatic post-deployment retraining
- Explainable artificial intelligence model to predict acute critical illness from electronic health records
- Towards computational fluorescence microscopy: Machine learning-based integrated prediction of morphological and molecular tumor profiles
- Explaining Convolutional Neural Networks using Softmax Gradient Layer-wise Relevance Propagation
- Model-agnostic explainable artificial intelligence for object detection in image data
- Towards Interpretable and Robust Hand Detection via Pixel-wise Prediction
- Explainable automatic industrial carbon footprint estimation from bank transaction classification using natural language processing
- Global Explanations of Neural Networks: Mapping the Landscape of Predictions
- Contrastive Explanations with Local Foil Trees
- Explanation-Guided Training for Cross-Domain Few-Shot Classification
- xERTE: Explainable Reasoning on Temporal Knowledge Graphs for Forecasting Future Links
- Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
- A Detailed Study of Interpretability of Deep Neural Network based Top Taggers
- The Clever Hans Effect in Anomaly Detection
- Exploiting auto-encoders and segmentation methods for middle-level explanations of image classification systems
- Disentangled Explanations of Neural Network Predictions by Finding Relevant Subspaces
- Breaking Batch Normalization for better explainability of Deep Neural Networks through Layer-wise Relevance Propagation
- GAMI-Net: An Explainable Neural Network based on Generalized Additive Models with Structured Interactions
- Analyzing Atomic Interactions in Molecules as Learned by Neural Networks
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural Networks
- Simulator-based explanation and debugging of hazard-triggering events in DNN-based safety-critical systems
- Deeply Explain CNN via Hierarchical Decomposition
- Occlusion Sensitivity Analysis with Augmentation Subspace Perturbation in Deep Feature Space
- An Introduction to Deep Visual Explanation
- Identifying the relevant dependencies of the neural network response on characteristics of the input space
- TSGB: Target-Selective Gradient Backprop for Probing CNN Visual Saliency
- Controlled abstention neural networks for identifying skillful predictions for classification problems
- Identifying Opportunities for Skillful Weather Prediction with Interpretable Neural Networks
- Strategy to Increase the Safety of a DNN-based Perception for HAD Systems
- Towards Symbolic XAI -- Explanation Through Human Understandable Logical Relationships Between Features
- Sparse Oblique Decision Trees: A Tool to Understand and Manipulate Neural Net Features
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Search for Higgs boson and observation of Z boson through their decay into a charm quark-antiquark pair in boosted topologies in proton-proton collisions at = 13 TeV
- What went wrong and when? Instance-wise Feature Importance for Time-series Models
- AdjointNet: Constraining machine learning models with physics-based codes
- TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
- LioNets: A Neural-Specific Local Interpretation Technique Exploiting Penultimate Layer Information
- Automatic explanation of the classification of Spanish legal judgments in jurisdiction-dependent law categories with tree estimators
- Relating Input Concepts to Convolutional Neural Network Decisions
- Trustworthy Convolutional Neural Networks: A Gradient Penalized-based Approach
- T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
- Interpreting Deep Neural Networks Through Variable Importance
- Understanding Individual Decisions of CNNs via Contrastive Backpropagation
- p-FP: Extraction, Classification, and Prediction of Website Fingerprints with Deep Learning
- A simple defense against adversarial attacks on heatmap explanations
- Interpretable Convolutional Neural Networks via Feedforward Design
- Debugging Tests for Model Explanations
- Interpretation of Deep Temporal Representations by Selective Visualization of Internally Activated Nodes
- Towards explainable artificial intelligence (XAI) for early anticipation of traffic accidents
- More cat than cute? Interpretable Prediction of Adjective-Noun Pairs
- Predicting the Travel Distance of Patients to Access Healthcare using Deep Neural Networks
- MeLIME: Meaningful Local Explanation for Machine Learning Models
- Why are Saliency Maps Noisy? Cause of and Solution to Noisy Saliency Maps
- XRAI: Better Attributions Through Regions
- Discriminating Spatial and Temporal Relevance in Deep Taylor Decompositions for Explainable Activity Recognition
- Reasoning on Knowledge Graphs with Debate Dynamics
- Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence Models
- InFIP: An Explainable DNN Intellectual Property Protection Method based on Intrinsic Features
- Improving Feature Attribution through Input-specific Network Pruning
- Robust Models Are More Interpretable Because Attributions Look Normal
- Evaluating the feasibility of interpretable machine learning for globular cluster detection
- Usefulness of interpretability methods to explain deep learning based plant stress phenotyping
- A Survey on Understanding, Visualizations, and Explanation of Deep Neural Networks
- Aggregating explanation methods for stable and robust explainability
- Engineering problems in machine learning systems
- AutoSourceID-FeatureExtractor. Optical image analysis using a two-step mean variance estimation network for feature estimation and uncertainty characterisation
- Interpreting Interpretations: Organizing Attribution Methods by Criteria
- A Neural Network Perturbation Theory Based on the Born Series
- Satellite galaxies' drag on field stars in the Milky Way
- Logic Programming and Machine Ethics
- On The Coherence of Quantitative Evaluation of Visual Explanations
- Explaining Motion Relevance for Activity Recognition in Video Deep Learning Models
- Multi-Scale Neural network for EEG Representation Learning in BCI
- Better Understanding Differences in Attribution Methods via Systematic Evaluations
- Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks
- Exploring explicit coarse-grained structure in artificial neural networks
- Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
- Understanding the Dependence of Perception Model Competency on Regions in an Image
- Explainable AI for ML jet taggers using expert variables and layerwise relevance propagation
- Challenges for cognitive decoding using deep learning methods
- On Quantitative Evaluations of Counterfactuals
- Rethinking Positive Aggregation and Propagation of Gradients in Gradient-based Saliency Methods
- An Empirical Study of Explainable AI Techniques on Deep Learning Models For Time Series Tasks
- Towards Robust Explanations for Deep Neural Networks
- Towards Interpretable Ensemble Learning for Image-based Malware Detection
- Explainable Product Search with a Dynamic Relation Embedding Model
- Evaluation, Tuning and Interpretation of Neural Networks for Meteorological Applications
- Sequential Explanations with Mental Model-Based Policies
- A psychophysics approach for quantitative comparison of interpretable computer vision models
- Born Identity Network: Multi-way Counterfactual Map Generation to Explain a Classifier's Decision
- Towards Privacy-preserving Explanations in Medical Image Analysis
- Achieving Explainability for Plant Disease Classification with Disentangled Variational Autoencoders
- Multi-Instance Multi-Scale CNN for Medical Image Classification
- Explain to Fix: A Framework to Interpret and Correct DNN Object Detector Predictions
- Efficient Image Evidence Analysis of CNN Classification Results
- Embedded Encoder-Decoder in Convolutional Networks Towards Explainable AI
- An Efficient Explorative Sampling Considering the Generative Boundaries of Deep Generative Neural Networks
- Quantitative and Qualitative Evaluation of Explainable Deep Learning Methods for Ophthalmic Diagnosis
- Between Homomorphic Signal Processing and Deep Neural Networks: Constructing Deep Algorithms for Polyphonic Music Transcription
- Gradient Weighted Superpixels for Interpretability in CNNs
- Interpreting deep learning-based stellar mass estimation via causal analysis and mutual information decomposition
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- Switched linear projections for neural network interpretability
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- SelfExplain: A Self-Explaining Architecture for Neural Text Classifiers
- Debate Dynamics for Human-comprehensible Fact-checking on Knowledge Graphs
- Explainability Guided Multi-Site COVID-19 CT Classification
- Explainable 3D Convolutional Neural Networks by Learning Temporal Transformations
- Improving Interpretability of Deep Neural Networks in Medical Diagnosis by Investigating the Individual Units
- Valid Explanations for Learning to Rank Models
- Neuron ranking -- an informed way to condense convolutional neural networks architecture
- How do Convolutional Neural Networks Learn Design?
- Evaluating the Reliability of Self-Explanations in Large Language Models
- Towards Deep Learning Models Resistant to Large Perturbations
- Data-Adaptive Discriminative Feature Localization with Statistically Guaranteed Interpretation
- Responsible LLM Deployment for High-Stake Decisions by Decentralized Technologies and Human-AI Interactions
- Heat and Blur: An Effective and Fast Defense Against Adversarial Examples
- Discriminative Attribution from Counterfactuals
- Debiasing Convolutional Neural Networks via Meta Orthogonalization
- A General Taylor Framework for Unifying and Revisiting Attribution Methods
- SPARK: Static Program Analysis Reasoning and Retrieving Knowledge
- Interpreting Deep Learning Model Using Rule-based Method
- Interpreting Undesirable Pixels for Image Classification on Black-Box Models
- A Tour of Convolutional Networks Guided by Linear Interpreters
- NeuroMask: Explaining Predictions of Deep Neural Networks through Mask Learning
- Explainability-aided Domain Generalization for Image Classification
- Femtosecond pulse parameter estimation from photoelectron momenta using machine learning
- Fast and Accurate Explanations of Distance-Based Classifiers by Uncovering Latent Explanatory Structures
- Weakly-Supervised Cell Tracking via Backward-and-Forward Propagation
- Explainable Adversarial Attacks on Coarse-to-Fine Classifiers
- Faster ISNet for Background Bias Mitigation on Deep Neural Networks
- Gradient Frequency Modulation for Visually Explaining Video Understanding Models
- Understanding Patch-Based Learning by Explaining Predictions
- A Categorisation of Post-hoc Explanations for Predictive Models
- Relevance Attack on Detectors
- Contrastive ACE: Domain Generalization Through Alignment of Causal Mechanisms
- Understanding of Kernels in CNN Models by Suppressing Irrelevant Visual Features in Images
- Layer-Wise Interpretation of Deep Neural Networks Using Identity Initialization
- Understanding the wiring evolution in differentiable neural architecture search
- Trustworthy Data-driven Chronological Age Estimation from Panoramic Dental Images
- Attribution Mask: Filtering Out Irrelevant Features By Recursively Focusing Attention on Inputs of DNNs
- Improving Attribution Methods by Learning Submodular Functions
- An Overview of Computational Approaches for Interpretation Analysis
- Interpreting Deep Neural Networks with Relative Sectional Propagation by Analyzing Comparative Gradients and Hostile Activations
- Efficient Modelling Across Time of Human Actions and Interactions
- Explain and Improve: LRP-Inference Fine-Tuning for Image Captioning Models
- Taylor saves for later: disentanglement for video prediction using Taylor representation
- Unsupervised Detection and Explanation of Latent-class Contextual Anomalies
- Deep Relevance Regularization: Interpretable and Robust Tumor Typing of Imaging Mass Spectrometry Data
- A Case Study of Deep-Learned Activations via Hand-Crafted Audio Features
- Learning Propagation Rules for Attribution Map Generation
- Cost-Sensitive Feature-Value Acquisition Using Feature Relevance
- Analysis of Video Feature Learning in Two-Stream CNNs on the Example of Zebrafish Swim Bout Classification
- How to Explain Neural Networks: an Approximation Perspective
- Understanding Convolutional Neural Networks with A Mathematical Model
- Explainable Deep Modeling of Tabular Data using TableGraphNet
- Sanity Simulations for Saliency Methods
- Self-learn to Explain Siamese Networks Robustly
- Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop
- Mutual Information Preserving Back-propagation: Learn to Invert for Faithful Attribution
- Towards Comparative Physical Interpretation of Spatial Variability Aware Neural Networks: A Summary of Results
- Weakly Supervised Recovery of Semantic Attributes
- Network Analysis for Explanation
- Explaining a prediction in some nonlinear models
- Interpreting deep urban sound classification using Layer-wise Relevance Propagation
- IMPACTX: improving model performance by appropriately constraining the training with teacher explanations
- Explaining Representation by Mutual Information
- On the Use of Interpretable Machine Learning for the Management of Data Quality
- Explainable Adversarial Attacks in Deep Neural Networks Using Activation Profiles