Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
arXiv:1605.01713
Abstract
Note: This paper describes an older version of DeepLIFT. See https://arxiv.org/abs/1704.02685 for the newer version. Original abstract follows: The purported "black box" nature of neural networks is a barrier to adoption in applications where interpretability is essential. Here we present DeepLIFT (Learning Important FeaTures), an efficient and effective method for computing importance scores in a neural network. DeepLIFT compares the activation of each neuron to its 'reference activation' and assigns contribution scores according to the difference. We apply DeepLIFT to models trained on natural images and genomic data, and show significant advantages over gradient-based methods.
6 pages, 3 figures, this is an older version; see https://arxiv.org/abs/1704.02685 for the newer version
References in corpus (2)
Cited by in corpus (139)
- A Unified Approach to Interpreting Model Predictions
- Axiomatic Attribution for Deep Networks
- Learning Important Features Through Propagating Activation Differences
- Interpretable machine learning: definitions, methods, and applications
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Understanding Black-box Predictions via Influence Functions
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- A Survey on the Explainability of Supervised Machine Learning
- Explaining Explanations in AI
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- Higher-Order Explanations of Graph Neural Networks via Relevant Walks
- Explainable AI for Trees: From Local Explanations to Global Understanding
- explAIner: A Visual Analytics Framework for Interactive and Explainable Machine Learning
- Guidelines and Evaluation of Clinical Explainable AI in Medical Image Analysis
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- Explainable Deep Reinforcement Learning: State of the Art and Challenges
- Feature relevance quantification in explainable AI: A causal problem
- Explaining by Removing: A Unified Framework for Model Explanation
- Neural Network Attribution Methods for Problems in Geoscience: A Novel Synthetic Benchmark Dataset
- An unexpected unity among methods for interpreting model predictions
- Hierarchical interpretations for neural network predictions
- Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience
- On the Robustness of Interpretability Methods
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- Building and Interpreting Deep Similarity Models
- Evaluating Explainable AI on a Multi-Modal Medical Imaging Task: Can Existing Algorithms Fulfill Clinical Requirements?
- On the (In)fidelity and Sensitivity for Explanations
- Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization
- AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech
- Explaining the Explainer: A First Theoretical Analysis of LIME
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Values Approximation
- Interpreting deep learning models for weak lensing
- Visualizing Deep Networks by Optimizing with Integrated Gradients
- Symbolic Execution for Deep Neural Networks
- When Explanations Lie: Why Many Modified BP Attributions Fail
- MAGIX: Model Agnostic Globally Interpretable Explanations
- Regularizing Black-box Models for Improved Interpretability
- Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
- RIDDLE: Race and ethnicity Imputation from Disease history with Deep LEarning
- DeepFaceLIFT: Interpretable Personalized Models for Automatic Estimation of Self-Reported Pain
- Shapley Flow: A Graph-based Approach to Interpreting Model Predictions
- Neural Machine Reading Comprehension: Methods and Trends
- DeepVATS: Deep Visual Analytics for Time Series
- Towards Best Practice of Interpreting Deep Learning Models for EEG-based Brain Computer Interfaces
- Weakly Supervised Deep Learning for COVID-19 Infection Detection and Classification from CT Images
- Interpreting the Predictions of Complex ML Models by Layer-wise Relevance Propagation
- deepTarget: End-to-end Learning Framework for microRNA Target Prediction using Deep Recurrent Neural Networks
- Privacy Meets Explainability: A Comprehensive Impact Benchmark
- Meta-trained agents implement Bayes-optimal agents
- Regularization Learning Networks: Deep Learning for Tabular Datasets
- I-SPLIT: Deep Network Interpretability for Split Computing
- Explainable Differential Privacy-Hyperdimensional Computing for Balancing Privacy and Transparency in Additive Manufacturing Monitoring
- Feature Removal Is a Unifying Principle for Model Explanation Methods
- Interpreting Super-Resolution Networks with Local Attribution Maps
- iSEA: An Interactive Pipeline for Semantic Error Analysis of NLP Models
- xGEMs: Generating Examplars to Explain Black-Box Models
- Best of both worlds: local and global explanations with human-understandable concepts
- Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks
- Individual Explanations in Machine Learning Models: A Survey for Practitioners
- A Gradient Mapping Guided Explainable Deep Neural Network for Extracapsular Extension Identification in 3D Head and Neck Cancer Computed Tomography Images
- Exemplars and Counterexemplars Explanations for Image Classifiers, Targeting Skin Lesion Labeling
- Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis
- Heterogeneous Graph Neural Networks with Post-hoc Explanations for Multi-modal and Explainable Land Use Inference
- Explainable Deep Reinforcement Learning for UAV Autonomous Navigation
- Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring
- LioNets: A Neural-Specific Local Interpretation Technique Exploiting Penultimate Layer Information
- Transformation Importance with Applications to Cosmology
- Interpreting Deep Neural Networks Through Variable Importance
- Improving the Transferability of Adversarial Attacks on Face Recognition with Beneficial Perturbation Feature Augmentation
- Towards a Unified Evaluation of Explanation Methods without Ground Truth
- Efficient Search for Diverse Coherent Explanations
- EX-RAY: Distinguishing Injected Backdoor from Natural Features in Neural Networks by Examining Differential Feature Symmetry
- My Teacher Thinks The World Is Flat! Interpreting Automatic Essay Scoring Mechanism
- Achievements and Challenges in Explaining Deep Learning based Computer-Aided Diagnosis Systems
- Debugging Tests for Model Explanations
- XRAI: Better Attributions Through Regions
- Whatcha lookin' at? DeepLIFTing BERT's Attention in Question Answering
- Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence Models
- Interpreting Vision and Language Generative Models with Semantic Visual Priors
- One Explanation is Not Enough: Structured Attention Graphs for Image Classification
- Shapley Value as Principled Metric for Structured Network Pruning
- Engineering problems in machine learning systems
- Explainable Multivariate Time Series Classification: A Deep Neural Network Which Learns To Attend To Important Variables As Well As Informative Time Intervals
- Can We Trust Your Explanations? Sanity Checks for Interpreters in Android Malware Analysis
- An Empirical Study of Explainable AI Techniques on Deep Learning Models For Time Series Tasks
- Inducing Causal Structure for Interpretable Neural Networks
- Passive Attention in Artificial Neural Networks Predicts Human Visual Selectivity
- Discretized Integrated Gradients for Explaining Language Models
- Individual Explanations in Machine Learning Models: A Case Study on Poverty Estimation
- A Review of Explainable Artificial Intelligence in Manufacturing
- GuidedMix-Net: Learning to Improve Pseudo Masks Using Labeled Images as Reference
- A Causal Lens for Peeking into Black Box Predictive Models: Predictive Model Interpretation via Causal Attribution
- SCOUT: Self-aware Discriminant Counterfactual Explanations
- Explaining AlphaGo: Interpreting Contextual Effects in Neural Networks
- GANMEX: One-vs-One Attributions Guided by GAN-based Counterfactual Explanation Baselines
- Assessment of the Reliablity of a Model's Decision by Generalizing Attribution to the Wavelet Domain
- Quantitative and Qualitative Evaluation of Explainable Deep Learning Methods for Ophthalmic Diagnosis
- Defining and Quantifying the Emergence of Sparse Concepts in DNNs
- AES Systems Are Both Overstable And Oversensitive: Explaining Why And Proposing Defenses
- Explainable Deep Image Classifiers for Skin Lesion Diagnosis
- Finding Discriminative Filters for Specific Degradations in Blind Super-Resolution
- iGOS++: Integrated Gradient Optimized Saliency by Bilateral Perturbations
- Paying Attention to Attention: Highlighting Influential Samples in Sequential Analysis
- Toward the Understanding of Deep Text Matching Models for Information Retrieval
- Explaining Regression Based Neural Network Model
- Shapley Explanation Networks
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- Evaluating Tree Explanation Methods for Anomaly Reasoning: A Case Study of SHAP TreeExplainer and TreeInterpreter
- Regularizing Black-box Models for Improved Interpretability (HILL 2019 Version)
- Evaluating the Use of Reconstruction Error for Novelty Localization
- The Generalizability of Explanations
- A General Taylor Framework for Unifying and Revisiting Attribution Methods
- On Spectral Properties of Gradient-based Explanation Methods
- Do Input Gradients Highlight Discriminative Features?
- Human-in-the-loop model explanation via verbatim boundary identification in generated neighborhoods
- Simplifying the explanation of deep neural networks with sufficient and necessary feature-sets: case of text classification
- Improved Feature Importance Computations for Tree Models: Shapley vs. Banzhaf
- Discriminative Attribution from Counterfactuals
- The Eval4NLP Shared Task on Explainable Quality Estimation: Overview and Results
- Bit Error Tolerance Metrics for Binarized Neural Networks
- Integrated Grad-CAM: Sensitivity-Aware Visual Explanation of Deep Convolutional Networks via Integrated Gradient-Based Scoring
- Deep Learning Based Decision Support for Medicine -- A Case Study on Skin Cancer Diagnosis
- Pattern-Guided Integrated Gradients
- Training Machine Learning Models by Regularizing their Explanations
- Noise Modulation: Let Your Model Interpret Itself
- Network Analysis for Explanation
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- Robustness of different loss functions and their impact on networks learning capability
- Visual Summary of Value-level Feature Attribution in Prediction Classes with Recurrent Neural Networks
- Learning Propagation Rules for Attribution Map Generation
- Streamlining models with explanations in the learning loop
- Evaluating Attribution Methods using White-Box LSTMs
- Deep Relevance Regularization: Interpretable and Robust Tumor Typing of Imaging Mass Spectrometry Data
- Causal Abstractions of Neural Networks
- Weakly Supervised Recovery of Semantic Attributes
- e-QRAQ: A Multi-turn Reasoning Dataset and Simulator with Explanations
- Rank Projection Trees for Multilevel Neural Network Interpretation