Learning Important Features Through Propagating Activation Differences
arXiv:1704.02685
Abstract
The purported "black box" nature of neural networks is a barrier to adoption in applications where interpretability is essential. Here we present DeepLIFT (Deep Learning Important FeaTures), a method for decomposing the output prediction of a neural network on a specific input by backpropagating the contributions of all neurons in the network to every feature of the input. DeepLIFT compares the activation of each neuron to its 'reference activation' and assigns contribution scores according to the difference. By optionally giving separate consideration to positive and negative contributions, DeepLIFT can also reveal dependencies which are missed by other approaches. Scores can be computed efficiently in a single backward pass. We apply DeepLIFT to models trained on MNIST and simulated genomic data, and show significant advantages over gradient-based methods. Video tutorial: http://goo.gl/qKb7pL, ICML slides: bit.ly/deeplifticmlslides, ICML talk: https://vimeo.com/238275076, code: http://goo.gl/RM8jvH.
Updated to include changes present in the ICML camera-ready paper, and other small corrections
References in corpus (5)
- Striving for Simplicity: The All Convolutional Net
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- An unexpected unity among methods for interpreting model predictions
- Investigating the influence of noise and distractors on the interpretation of neural networks
- Gradients of Counterfactuals
Cited by in corpus (449)
- A Unified Approach to Interpreting Model Predictions
- Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- On the Opportunities and Risks of Foundation Models
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Explaining Explanations in AI
- GNNExplainer: Generating Explanations for Graph Neural Networks
- Captum: A unified and generic model interpretability library for PyTorch
- Towards Explainable Artificial Intelligence
- Learning data driven discretizations for partial differential equations
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- Eigen-CAM: Class Activation Map using Principal Components
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Explanations in Autonomous Driving: A Survey
- Explaining a Series of Models by Propagating Shapley Values
- explAIner: A Visual Analytics Framework for Interactive and Explainable Machine Learning
- Polyconvex anisotropic hyperelasticity with neural networks
- Deep Learning in Pharmacogenomics: From Gene Regulation to Patient Stratification
- Technical Note on Transcription Factor Motif Discovery from Importance Scores (TF-MoDISco) version 0.5.6.5
- Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models
- Evaluating explainable artificial intelligence methods for multi-label deep learning classification tasks in remote sensing
- Enhancing Discrete Choice Models with Representation Learning
- XCM: An Explainable Convolutional Neural Network for Multivariate Time Series Classification
- Automatic Sleep Staging of EEG Signals: Recent Development, Challenges, and Future Directions
- Explanations can be manipulated and geometry is to blame
- Interpreting Blackbox Models via Model Extraction
- Evaluation of post-hoc interpretability methods in time-series classification
- Post-hoc explanation of black-box classifiers using confident itemsets
- On the Explainability of Natural Language Processing Deep Models
- Fooling Neural Network Interpretations via Adversarial Model Manipulation
- Machine learning and AI-based approaches for bioactive ligand discovery and GPCR-ligand recognition
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications
- From Clustering to Cluster Explanations via Neural Networks
- Robust Explainability: A Tutorial on Gradient-Based Attribution Methods for Deep Neural Networks
- Learning to Explain: An Information-Theoretic Perspective on Model Interpretation
- Neural Network Attribution Methods for Problems in Geoscience: A Novel Synthetic Benchmark Dataset
- Towards Explaining Anomalies: A Deep Taylor Decomposition of One-Class Models
- TimeSHAP: Explaining Recurrent Models through Sequence Perturbations
- TSViz: Demystification of Deep Learning Models for Time-Series Analysis
- ISeeU: Visually interpretable deep learning for mortality prediction inside the ICU
- Exploring Interpretable LSTM Neural Networks over Multi-Variable Data
- Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks
- Feature selection revisited in the single-cell era
- Explainable Artificial Intelligence: a Systematic Review
- How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations
- To trust or not to trust an explanation: using LEAF to evaluate local linear XAI methods
- Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- Explainable Diabetic Retinopathy Detection and Retinal Image Generation
- Explaining Anomalies Detected by Autoencoders Using SHAP
- Interpretable deep learning for guided structure-property explorations in photovoltaics
- Deep neural network improves the estimation of polygenic risk scores for breast cancer
- A general framework for inference on algorithm-agnostic variable importance
- XAIR: A Framework of Explainable AI in Augmented Reality
- An analysis on the use of autoencoders for representation learning: fundamentals, learning task case studies, explainability and challenges
- Interpretation of Neural Networks is Fragile
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- Explainable Machine Learning for Public Policy: Use Cases, Gaps, and Research Directions
- Counterfactual State Explanations for Reinforcement Learning Agents via Generative Deep Learning
- Deep Learning Based Cloud Cover Parameterization for ICON
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Evaluating Explainable AI on a Multi-Modal Medical Imaging Task: Can Existing Algorithms Fulfill Clinical Requirements?
- Neuron Shapley: Discovering the Responsible Neurons
- On the (In)fidelity and Sensitivity for Explanations
- Security and Privacy Issues in Deep Learning
- Explainability: Relevance based Dynamic Deep Learning Algorithm for Fault Detection and Diagnosis in Chemical Processes
- Explainability Techniques for Graph Convolutional Networks
- A Survey of Deep Learning for Scientific Discovery
- DIG: A Turnkey Library for Diving into Graph Deep Learning Research
- Testing and verification of neural-network-based safety-critical control software: A systematic literature review
- Real-time Neural Network Inference on Extremely Weak Devices: Agile Offloading with Explainable AI
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- A Theoretical Explanation for Perplexing Behaviors of Backpropagation-based Visualizations
- Restricting the Flow: Information Bottlenecks for Attribution
- Medical Imaging and Machine Learning
- Explaining the Explainer: A First Theoretical Analysis of LIME
- SS-CAM: Smoothed Score-CAM for Sharper Visual Feature Localization
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Beyond saliency: understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation
- Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Values Approximation
- Towards Evaluating Explanations of Vision Transformers for Medical Imaging
- Parkinson's Disease Recognition Using SPECT Image and Interpretable AI: A Tutorial
- Causally-informed deep learning to improve climate models and projections
- RRWaveNet: A Compact End-to-End Multi-Scale Residual CNN for Robust PPG Respiratory Rate Estimation
- When Explanations Lie: Why Many Modified BP Attributions Fail
- PGM-Explainer: Probabilistic Graphical Model Explanations for Graph Neural Networks
- Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
- Towards Better Understanding Attribution Methods
- Interpreting Graph Neural Networks for NLP With Differentiable Edge Masking
- NormLime: A New Feature Importance Metric for Explaining Deep Neural Networks
- Shedding Light on Black Box Machine Learning Algorithms: Development of an Axiomatic Framework to Assess the Quality of Methods that Explain Individual Predictions
- Do Not Trust Additive Explanations
- Can I Trust the Explainer? Verifying Post-hoc Explanatory Methods
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainability
- How does this interaction affect me? Interpretable attribution for feature interactions
- Dropout Feature Ranking for Deep Learning Models
- RIDDLE: Race and ethnicity Imputation from Disease history with Deep LEarning
- Considerations When Learning Additive Explanations for Black-Box Models
- A new interpretable unsupervised anomaly detection method based on residual explanation
- Neural Network Attributions: A Causal Perspective
- Interpreting Black Box Models via Hypothesis Testing
- Epistemic values in feature importance methods: Lessons from feminist epistemology
- Making Neural Networks Interpretable with Attribution: Application to Implicit Signals Prediction
- Autoencoder Node Saliency: Selecting Relevant Latent Representations
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
- Granger-causal Attentive Mixtures of Experts: Learning Important Features with Neural Networks
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
- A Survey on Neural Network Interpretability
- Understanding Integrated Gradients with SmoothTaylor for Deep Neural Network Attribution
- White Box Methods for Explanations of Convolutional Neural Networks in Image Classification Tasks
- Exploration of Interpretability Techniques for Deep COVID-19 Classification using Chest X-ray Images
- On the Art and Science of Machine Learning Explanations
- Towards falsifiable interpretability research
- Shapley Flow: A Graph-based Approach to Interpreting Model Predictions
- Global Aggregations of Local Explanations for Black Box models
- Inseq: An Interpretability Toolkit for Sequence Generation Models
- SHAP values for Explaining CNN-based Text Classification Models
- Counterfactual Explanation Based on Gradual Construction for Deep Networks
- Explaining Clinical Decision Support Systems in Medical Imaging using Cycle-Consistent Activation Maximization
- Regional Multi-scale Approach for Visually Pleasing Explanations of Deep Neural Networks
- Interpretability Beyond Classification Output: Semantic Bottleneck Networks
- Explaining Time Series Predictions with Dynamic Masks
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
- Bridging Adversarial Robustness and Gradient Interpretability
- Smoothed Geometry for Robust Attribution
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Self-Explaining Structures Improve NLP Models
- Investigating the significance of adversarial attacks and their relation to interpretability for radar-based human activity recognition systems
- Towards Frequency-Based Explanation for Robust CNN
- Infusing domain knowledge in AI-based "black box" models for better explainability with application in bankruptcy prediction
- Self-Attention Attribution: Interpreting Information Interactions Inside Transformer
- Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation
- Explaining Deep Neural Networks
- Embedding Deep Networks into Visual Explanations
- Fast Hierarchical Games for Image Explanations
- Saliency Methods for Explaining Adversarial Attacks
- Explaining Convolutional Neural Networks using Softmax Gradient Layer-wise Relevance Propagation
- Coalitional strategies for efficient individual prediction explanation
- Unwrapping The Black Box of Deep ReLU Networks: Interpretability, Diagnostics, and Simplification
- AI for Explaining Decisions in Multi-Agent Environments
- When Explainability Meets Adversarial Learning: Detecting Adversarial Examples using SHAP Signatures
- Predicting drug response of tumors from integrated genomic profiles by deep neural networks
- Global Explanations of Neural Networks: Mapping the Landscape of Predictions
- Interpret Federated Learning with Shapley Values
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
- Realised Volatility Forecasting: Machine Learning via Financial Word Embedding
- Explanation-Guided Training for Cross-Domain Few-Shot Classification
- DISCOVER: 2-D Multiview Summarization of Optical Coherence Tomography Angiography for Automatic Diabetic Retinopathy Diagnosis
- Evaluating the Correctness of Explainable AI Algorithms for Classification
- Deep Learning in Protein Structural Modeling and Design
- Adversarial Infidelity Learning for Model Interpretation
- Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based Explanation
- FairCanary: Rapid Continuous Explainable Fairness
- Refining Language Models with Compositional Explanations
- Interpreting Super-Resolution Networks with Local Attribution Maps
- Interpretable Deep Learning under Fire
- Certifiably Robust Interpretation in Deep Learning
- Breaking Batch Normalization for better explainability of Deep Neural Networks through Layer-wise Relevance Propagation
- On Baselines for Local Feature Attributions
- Mutual Information for Explainable Deep Learning of Multiscale Systems
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- A Rate-Distortion Framework for Explaining Neural Network Decisions
- Generative causal explanations of black-box classifiers
- Unbox the Black-box for the Medical Explainable AI via Multi-modal and Multi-centre Data Fusion: A Mini-Review, Two Showcases and Beyond
- Deeply Explain CNN via Hierarchical Decomposition
- Understanding Bias in Machine Learning
- On Explaining Your Explanations of BERT: An Empirical Study with Sequence Classification
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural Networks
- ML-LOO: Detecting Adversarial Examples with Feature Attribution
- Weakly-Supervised Action Localization and Action Recognition using Global-Local Attention of 3D CNN
- Domain Knowledge Aided Explainable Artificial Intelligence for Intrusion Detection and Response
- Human Attention in Fine-grained Classification
- Best of both worlds: local and global explanations with human-understandable concepts
- Interpretation of multi-label classification models using shapley values
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- Latent-CF: A Simple Baseline for Reverse Counterfactual Explanations
- Fooling Partial Dependence via Data Poisoning
- Feature construction using explanations of individual predictions
- Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation
- Learn to Interpret Atari Agents
- Benchmarking the Influence of Pre-training on Explanation Performance in MR Image Classification
- GeCo: Quality Counterfactual Explanations in Real Time
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- Joint Explainability and Sensitivity-Aware Federated Deep Learning for Transparent 6G RAN Slicing
- Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
- What is Interpretable? Using Machine Learning to Design Interpretable Decision-Support Systems
- Melody: Generating and Visualizing Machine Learning Model Summary to Understand Data and Classifiers Together
- Unveiling The Factors of Aesthetic Preferences with Explainable AI
- Shapley explainability on the data manifold
- Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis
- Towards Explainable Deep Learning for Credit Lending: A Case Study
- Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
- VizADS-B: Analyzing Sequences of ADS-B Images Using Explainable Convolutional LSTM Encoder-Decoder to Detect Cyber Attacks
- Bias, Fairness, and Accountability with AI and ML Algorithms
- ECINN: Efficient Counterfactuals from Invertible Neural Networks
- The Twin-System Approach as One Generic Solution for XAI: An Overview of ANN-CBR Twins for Explaining Deep Learning
- Explainable Deep Reinforcement Learning for UAV Autonomous Navigation
- Certification of embedded systems based on Machine Learning: A survey
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Fast Real-time Counterfactual Explanations
- Explaining in Style: Training a GAN to explain a classifier in StyleSpace
- Explaining Deep Neural Networks using Unsupervised Clustering
- A Note about: Local Explanation Methods for Deep Neural Networks lack Sensitivity to Parameter Values
- What went wrong and when? Instance-wise Feature Importance for Time-series Models
- Spatio-Temporal Perturbations for Video Attribution
- Sparse Oblique Decision Trees: A Tool to Understand and Manipulate Neural Net Features
- LioNets: A Neural-Specific Local Interpretation Technique Exploiting Penultimate Layer Information
- Model Explainability in Deep Learning Based Natural Language Processing
- Understanding Individual Decisions of CNNs via Contrastive Backpropagation
- Self-explanatory Deep Salient Object Detection
- Rationalizing Predictions by Adversarial Information Calibration
- SUBPLEX: Towards a Better Understanding of Black Box Model Explanations at the Subpopulation Level
- Multi-modal Machine Learning for Vehicle Rating Predictions Using Image, Text, and Parametric Data
- A principle feature analysis
- Interpreting Deep Neural Networks Through Variable Importance
- i-Align: an interpretable knowledge graph alignment model
- Proactive Pseudo-Intervention: Causally Informed Contrastive Learning For Interpretable Vision Models
- A Conceptual Framework for Establishing Trust in Real World Intelligent Systems
- Incorporating Priors with Feature Attribution on Text Classification
- A Baseline for Shapley Values in MLPs: from Missingness to Neutrality
- Computationally Efficient Feature Significance and Importance for Machine Learning Models
- A simple defense against adversarial attacks on heatmap explanations
- What evidence does deep learning model use to classify Skin Lesions?
- On the Evaluation of the Plausibility and Faithfulness of Sentiment Analysis Explanations
- Evaluating neural network explanation methods using hybrid documents and morphological agreement
- Rethinking the Role of Gradient-Based Attribution Methods for Model Interpretability
- Spatio-Temporal Momentum: Jointly Learning Time-Series and Cross-Sectional Strategies
- Identifying Dominant Industrial Sectors in Market States of the S&P 500 Financial Data
- Semantics for Global and Local Interpretation of Deep Neural Networks
- CAUSE: Learning Granger Causality from Event Sequences using Attribution Methods
- Order in the Court: Explainable AI Methods Prone to Disagreement
- Why are Saliency Maps Noisy? Cause of and Solution to Noisy Saliency Maps
- The Shapley Taylor Interaction Index
- Human-interpretable model explainability on high-dimensional data
- ProtoryNet - Interpretable Text Classification Via Prototype Trajectories
- Have We Learned to Explain?: How Interpretability Methods Can Learn to Encode Predictions in their Interpretations
- XRAI: Better Attributions Through Regions
- Interpretation of Deep Temporal Representations by Selective Visualization of Internally Activated Nodes
- Hide-and-Seek: A Template for Explainable AI
- Gaussian Mixture Models for Blended Photometric Redshifts
- Series Saliency: Temporal Interpretation for Multivariate Time Series Forecasting
- Robust Models Are More Interpretable Because Attributions Look Normal
- Developing Future Human-Centered Smart Cities: Critical Analysis of Smart City Security, Interpretability, and Ethical Challenges
- Information-Theoretic Visual Explanation for Black-Box Classifiers
- Whatcha lookin' at? DeepLIFTing BERT's Attention in Question Answering
- Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence Models
- Attention Mechanism for Multivariate Time Series Recurrent Model Interpretability Applied to the Ironmaking Industry
- Interpretable Graph Capsule Networks for Object Recognition
- Maximally Invariant Data Perturbation as Explanation
- Consistent Feature Selection for Analytic Deep Neural Networks
- Shapley Value as Principled Metric for Structured Network Pruning
- Improving Feature Attribution through Input-specific Network Pruning
- Identifying Galaxy Cluster Mergers with Deep Neural Networks using Idealized Compton-y and X-ray maps
- Modelling urban networks using Variational Autoencoders
- Optimising for Interpretability: Convolutional Dynamic Alignment Networks
- MonoNet: Towards Interpretable Models by Learning Monotonic Features
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations
- Distilling Ensemble of Explanations for Weakly-Supervised Pre-Training of Image Segmentation Models
- Explainable Multivariate Time Series Classification: A Deep Neural Network Which Learns To Attend To Important Variables As Well As Informative Time Intervals
- xAI-GAN: Enhancing Generative Adversarial Networks via Explainable AI Systems
- The Need for Standardized Explainability
- Aggregating explanation methods for stable and robust explainability
- Towards interpreting ML-based automated malware detection models: a survey
- Learning Deep Attribution Priors Based On Prior Knowledge
- Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity
- Utilizing Explainable AI for Quantization and Pruning of Deep Neural Networks
- X-SHAP: towards multiplicative explainability of Machine Learning
- In-Distribution Interpretability for Challenging Modalities
- Neural Architecture Search for Joint Optimization of Predictive Power and Biological Knowledge
- Interpretable Learning-to-Rank with Generalized Additive Models
- On The Coherence of Quantitative Evaluation of Visual Explanations
- Towards Human-Interpretable Prototypes for Visual Assessment of Image Classification Models
- Contextual Prediction Difference Analysis for Explaining Individual Image Classifications
- Convolutional Neural Network Interpretability with General Pattern Theory
- Unsupervised Detection of Distinctive Regions on 3D Shapes
- Just in Time: Personal Temporal Insights for Altering Model Decisions
- Human Understandable Explanation Extraction for Black-box Classification Models Based on Matrix Factorization
- Multi-Stage Influence Function
- On Iterative Neural Network Pruning, Reinitialization, and the Similarity of Masks
- Discretized Integrated Gradients for Explaining Language Models
- Framing Algorithmic Recourse for Anomaly Detection
- Investigating Saturation Effects in Integrated Gradients
- An Empirical Study of Explainable AI Techniques on Deep Learning Models For Time Series Tasks
- Knowledge-based XAI through CBR: There is more to explanations than models can tell
- Towards Robust Explanations for Deep Neural Networks
- Predicting traffic overflows on private peering
- Towards Interpretable Ensemble Learning for Image-based Malware Detection
- Quantitative Evaluation of Explainable Graph Neural Networks for Molecular Property Prediction
- Training and Predicting Visual Error for Real-Time Applications
- EMAP: Explanation by Minimal Adversarial Perturbation
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Towards Visually Explaining Video Understanding Networks with Perturbation
- Born Identity Network: Multi-way Counterfactual Map Generation to Explain a Classifier's Decision
- Attribution Analysis of Grammatical Dependencies in LSTMs
- Unveiling Black-boxes: Explainable Deep Learning Models for Patent Classification
- WSAM: Visual Explanations from Style Augmentation as Adversarial Attacker and Their Influence in Image Classification
- Identifying Critical Neurons in ANN Architectures using Mixed Integer Programming
- Gradient Weighted Superpixels for Interpretability in CNNs
- SCOUT: Self-aware Discriminant Counterfactual Explanations
- Technologies for Trustworthy Machine Learning: A Survey in a Socio-Technical Context
- Good-looking but Lacking Faithfulness: Understanding Local Explanation Methods through Trend-based Testing
- Neural Greedy Pursuit for Feature Selection
- GANMEX: One-vs-One Attributions Guided by GAN-based Counterfactual Explanation Baselines
- The Penalty Imposed by Ablated Data Augmentation
- A Causal Lens for Peeking into Black Box Predictive Models: Predictive Model Interpretation via Causal Attribution
- FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging
- A Spatiotemporal Radar-Based Precipitation Model for Water Level Prediction and Flood Forecasting
- Switched linear projections for neural network interpretability
- Designing Counterfactual Generators using Deep Model Inversion
- Many Faces of Feature Importance: Comparing Built-in and Post-hoc Feature Importance in Text Classification
- Evaluation of Saliency-based Explainability Method
- Optimal Piecewise Local-Linear Approximations
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- Improving Interpretability of Deep Neural Networks in Medical Diagnosis by Investigating the Individual Units
- Toward the Understanding of Deep Text Matching Models for Information Retrieval
- A Step Towards Exposing Bias in Trained Deep Convolutional Neural Network Models
- Towards Explanation of DNN-based Prediction with Guided Feature Inversion
- SelfExplain: A Self-Explaining Architecture for Neural Text Classifiers
- Scalable, Axiomatic Explanations of Deep Alzheimer's Diagnosis from Heterogeneous Data
- Model Interpretation and Explainability: Towards Creating Transparency in Prediction Models
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- What Do Adversarially Robust Models Look At?
- Explaining Image Classifiers using Statistical Fault Localization
- Neural Image Compression and Explanation
- Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding
- Evaluating Tree Explanation Methods for Anomaly Reasoning: A Case Study of SHAP TreeExplainer and TreeInterpreter
- FIND: Human-in-the-Loop Debugging Deep Text Classifiers
- Understanding and Diagnosing Vulnerability under Adversarial Attacks
- Fine Grained Dataflow Tracking with Proximal Gradients
- Partially Interpretable Estimators (PIE): Black-Box-Refined Interpretable Machine Learning
- CXR-Net: An Artificial Intelligence Pipeline for Quick Covid-19 Screening of Chest X-Rays
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Explainability Guided Multi-Site COVID-19 CT Classification
- Ontology-based Interpretable Machine Learning for Textual Data
- Leveraging Activation Maximization and Generative Adversarial Training to Recognize and Explain Patterns in Natural Areas in Satellite Imagery
- Measuring and improving the quality of visual explanations
- Shapley Explanation Networks
- MMLSpark: Unifying Machine Learning Ecosystems at Massive Scales
- FastSHAP: Real-Time Shapley Value Estimation
- Towards Interpretable and Transferable Speech Emotion Recognition: Latent Representation Based Analysis of Features, Methods and Corpora
- VBridge: Connecting the Dots Between Features and Data to Explain Healthcare Models
- A Comparison of Code Embeddings and Beyond
- Visualizing Color-wise Saliency of Black-Box Image Classification Models
- On the Robustness of Pretraining and Self-Supervision for a Deep Learning-based Analysis of Diabetic Retinopathy
- Class Introspection: A Novel Technique for Detecting Unlabeled Subclasses by Leveraging Classifier Explainability Methods
- Exploring layerwise decision making in DNNs
- Improved Feature Importance Computations for Tree Models: Shapley vs. Banzhaf
- Minimal Sufficient Views: A DNN model making predictions with more evidence has higher accuracy
- Harnessing value from data science in business: ensuring explainability and fairness of solutions
- Understanding Character Recognition using Visual Explanations Derived from the Human Visual System and Deep Networks
- Learning by Active Forgetting for Neural Networks
- Explaining GNN over Evolving Graphs using Information Flow
- Making Document-Level Information Extraction Right for the Right Reasons
- Skin Deep Unlearning: Artefact and Instrument Debiasing in the Context of Melanoma Classification
- Logic Traps in Evaluating Attribution Scores
- BERT-Beta: A Proactive Probabilistic Approach to Text Moderation
- Discriminative Attribution from Counterfactuals
- Uncertainty-aware INVASE: Enhanced Breast Cancer Diagnosis Feature Selection
- A General Taylor Framework for Unifying and Revisiting Attribution Methods
- The Eval4NLP Shared Task on Explainable Quality Estimation: Overview and Results
- Perceptual Score: What Data Modalities Does Your Model Perceive?
- Explanations for Occluded Images
- Do Input Gradients Highlight Discriminative Features?
- High resolution weakly supervised localization architectures for medical images
- Relevance Attack on Detectors
- Explainability-aided Domain Generalization for Image Classification
- DEPARA: Deep Attribution Graph for Deep Knowledge Transferability
- Towards Interpretable ANNs: An Exact Transformation to Multi-Class Multivariate Decision Trees
- Interpreting A Pre-trained Model Is A Key For Model Architecture Optimization: A Case Study On Wav2Vec 2.0
- DANCE: Enhancing saliency maps using decoys
- Human-grounded Evaluations of Explanation Methods for Text Classification
- A Tour of Convolutional Networks Guided by Linear Interpreters
- Deep interpretability for GWAS
- DRR4Covid: Learning Automated COVID-19 Infection Segmentation from Digitally Reconstructed Radiographs
- Evaluation of Local Explanation Methods for Multivariate Time Series Forecasting
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Explicating feature contribution using Random Forest proximity distances
- BayesGrad: Explaining Predictions of Graph Convolutional Networks
- Human-in-the-loop model explanation via verbatim boundary identification in generated neighborhoods
- Mutual Information Preserving Back-propagation: Learn to Invert for Faithful Attribution
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- Dividing Deep Learning Model for Continuous Anomaly Detection of Inconsistent ICT Systems
- CACTUS: Detecting and Resolving Conflicts in Objective Functions
- Human-Understandable Decision Making for Visual Recognition
- Attribution Mask: Filtering Out Irrelevant Features By Recursively Focusing Attention on Inputs of DNNs
- Diagnosis and Analysis of Celiac Disease and Environmental Enteropathy on Biopsy Images using Deep Learning Approaches
- XAlgo: a Design Probe of Explaining Algorithms' Internal States via Question-Answering
- Explaining a prediction in some nonlinear models
- Interpreting Deep Neural Networks with Relative Sectional Propagation by Analyzing Comparative Gradients and Hostile Activations
- Learning Propagation Rules for Attribution Map Generation
- Visualizing Classification Structure of Large-Scale Classifiers
- Marginal Contribution Feature Importance -- an Axiomatic Approach for The Natural Case
- Model-Agnostic Explanations using Minimal Forcing Subsets
- Play Fair: Frame Attributions in Video Models
- An Investigation of Language Model Interpretability via Sentence Editing
- AI challenges for predicting the impact of mutations on protein stability
- Defense Against Explanation Manipulation
- A Methodology for Exploring Deep Convolutional Features in Relation to Hand-Crafted Features with an Application to Music Audio Modeling
- Coalitional Bayesian Autoencoders -- Towards explainable unsupervised deep learning
- NeuroView: Explainable Deep Network Decision Making
- Comparative Study of Language Models on Cross-Domain Data with Model Agnostic Explainability
- An Empirical Study towards Understanding How Deep Convolutional Nets Recognize Falls
- Logic Constraints to Feature Importances
- Prediction of Hereditary Cancers Using Neural Networks
- Neural Supervised Domain Adaptation by Augmenting Pre-trained Models with Random Units
- Improving Attribution Methods by Learning Submodular Functions
- A Framework for Rationale Extraction for Deep QA models
- Rank Projection Trees for Multilevel Neural Network Interpretation
- How to Explain Neural Networks: an Approximation Perspective
- Sanity Simulations for Saliency Methods
- Regularizing Explanations in Bayesian Convolutional Neural Networks
- Information-theoretic Evolution of Model Agnostic Global Explanations
- Efficient Modelling Across Time of Human Actions and Interactions
- Local Explanation of Dialogue Response Generation
- SEEN: Sharpening Explanations for Graph Neural Networks using Explanations from Neighborhoods
- Explainability Requires Interactivity
- Understanding Misclassifications by Attributes
- Perturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
- Less is More: Feature Selection for Adversarial Robustness with Compressive Counter-Adversarial Attacks
- A Robust Unsupervised Ensemble of Feature-Based Explanations using Restricted Boltzmann Machines
- Counterfactual Graphs for Explainable Classification of Brain Networks
- T3-Vis: a visual analytic framework for Training and fine-Tuning Transformers in NLP
- Unsupervised Detection and Explanation of Latent-class Contextual Anomalies
- Thermostat: A Large Collection of NLP Model Explanations and Analysis Tools
- Towards Self-Explainable Graph Neural Network
- Longitudinal Distance: Towards Accountable Instance Attribution
- Saliency strikes back: How filtering out high frequencies improves white-box explanations
- Leveraging Model Interpretability and Stability to increase Model Robustness
- Interpretable Summaries of Black Box Incident Triaging with Subgroup Discovery
- Unsupervised discovery of Interpretable Visual Concepts
- Regularized Operating Envelope with Interpretability and Implementability Constraints
- Don't be fooled: label leakage in explanation methods and the importance of their quantitative evaluation
- Counterfactual Explanations via Latent Space Projection and Interpolation
- Interpretability of Blackbox Machine Learning Models through Dataview Extraction and Shadow Model creation
- Biophysical models of cis-regulation as interpretable neural networks
- Explainable Deep Reinforcement Learning for Portfolio Management: An Empirical Approach
- Explainable Deep Modeling of Tabular Data using TableGraphNet
- Balancing Robustness and Sensitivity using Feature Contrastive Learning
- Probabilistic Selective Encryption of Convolutional Neural Networks for Hierarchical Services
- Based on Graph-VAE Model to Predict Student's Score
- Do not explain without context: addressing the blind spot of model explanations
- Logic and the -Simplicial Transformer
- A Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without Segmentation
- Automated Dependence Plots
- Rigorous Explanation of Inference on Probabilistic Graphical Models
- Visual Summary of Value-level Feature Attribution in Prediction Classes with Recurrent Neural Networks
- Explain and Improve: LRP-Inference Fine-Tuning for Image Captioning Models