Axiomatic Attribution for Deep Networks
arXiv:1703.01365
Abstract
We study the problem of attributing the prediction of a deep network to its input features, a problem previously studied by several other works. We identify two fundamental axioms---Sensitivity and Implementation Invariance that attribution methods ought to satisfy. We show that they are not satisfied by most known attribution methods, which we consider to be a fundamental weakness of those methods. We use the axioms to guide the design of a new attribution method called Integrated Gradients. Our method requires no modification to the original network and is extremely simple to implement; it just needs a few calls to the standard gradient operator. We apply this method to a couple of image models, a couple of text models and a chemistry model, demonstrating its ability to debug networks, to extract rules from a network, and to enable users to engage with models better.
References in corpus (7)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Striving for Simplicity: The All Convolutional Net
- Learning Important Features Through Propagating Activation Differences
- Understanding Neural Networks Through Deep Visualization
- Going Deeper with Convolutions
- Convolutional Neural Networks for Sentence Classification
- An unexpected unity among methods for interpreting model predictions
Cited by in corpus (664)
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- On the Opportunities and Risks of Foundation Models
- Interpretable machine learning: definitions, methods, and applications
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies
- GNNExplainer: Generating Explanations for Graph Neural Networks
- Captum: A unified and generic model interpretability library for PyTorch
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Towards Explainable Artificial Intelligence
- This Looks Like That: Deep Learning for Interpretable Image Recognition
- Learning data driven discretizations for partial differential equations
- The What-If Tool: Interactive Probing of Machine Learning Models
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- Attention is not Explanation
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- Analyzing Federated Learning through an Adversarial Lens
- Fast Parallel Hypertree Decompositions in Logarithmic Recursion Depth
- Contrastive Learning of Subject-Invariant EEG Representations for Cross-Subject Emotion Recognition
- Explaining a Series of Models by Propagating Shapley Values
- explAIner: A Visual Analytics Framework for Interactive and Explainable Machine Learning
- What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use
- Pathologies of Neural Models Make Interpretations Difficult
- Towards Out-Of-Distribution Generalization: A Survey
- A Survey of the State of Explainable AI for Natural Language Processing
- Understanding Global Feature Contributions With Additive Importance Measures
- Post-hoc Interpretability for Neural NLP: A Survey
- A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual & Group Unfairness via Inequality Indices
- Technical Note on Transcription Factor Motif Discovery from Importance Scores (TF-MoDISco) version 0.5.6.5
- GAN Dissection: Visualizing and Understanding Generative Adversarial Networks
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- Evaluating explainable artificial intelligence methods for multi-label deep learning classification tasks in remote sensing
- Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models
- Feature relevance quantification in explainable AI: A causal problem
- Automatic Sleep Staging of EEG Signals: Recent Development, Challenges, and Future Directions
- Explanations can be manipulated and geometry is to blame
- NeuralHydrology -- Interpreting LSTMs in Hydrology
- TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
- Using Attribution to Decode Dataset Bias in Neural Network Models for Chemistry
- Explaining by Removing: A Unified Framework for Model Explanation
- Levels of explainable artificial intelligence for human-aligned conversational explanations
- Fooling Neural Network Interpretations via Adversarial Model Manipulation
- Revisiting Deep Learning Models for Tabular Data
- Toward Explainable AI for Regression Models
- If I Hear You Correctly: Building and Evaluating Interview Chatbots with Active Listening Skills
- From Clustering to Cluster Explanations via Neural Networks
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications
- Robust Explainability: A Tutorial on Gradient-Based Attribution Methods for Deep Neural Networks
- Rethinking CNN Models for Audio Classification
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- Neural Network Attribution Methods for Problems in Geoscience: A Novel Synthetic Benchmark Dataset
- Explainable Deep One-Class Classification
- RUDDER: Return Decomposition for Delayed Rewards
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Adversarial Examples on Graph Data: Deep Insights into Attack and Defense
- Towards Explaining Anomalies: A Deep Taylor Decomposition of One-Class Models
- Acquisition of Chess Knowledge in AlphaZero
- Impossibility Theorems for Feature Attribution
- Detecting hidden signs of diabetes in external eye photographs
- Multi-Objective Molecule Generation using Interpretable Substructures
- Hierarchical interpretations for neural network predictions
- Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks
- Rethinking Search: Making Domain Experts out of Dilettantes
- Learning and Evaluating Representations for Deep One-class Classification
- On the Generalizability of Neural Program Models with respect to Semantic-Preserving Program Transformations
- Explainable Artificial Intelligence: a Systematic Review
- To trust or not to trust an explanation: using LEAF to evaluate local linear XAI methods
- Benchmarking Attribution Methods with Relative Feature Importance
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience
- Towards Automatic Concept-based Explanations
- Explainable Neural Networks based on Additive Index Models
- DeepRobust: A PyTorch Library for Adversarial Attacks and Defenses
- Machine learning-assisted design of material properties
- Semantic Robustness of Models of Source Code
- Explainable Artificial Intelligence for Process Mining: A General Overview and Application of a Novel Local Explanation Approach for Predictive Process Monitoring
- A general framework for inference on algorithm-agnostic variable importance
- On the Robustness of Interpretability Methods
- A Survey of Explainable Artificial Intelligence (XAI) in Financial Time Series Forecasting
- Machine Learning in a data-limited regime: Augmenting experiments with synthetic data uncovers order in crumpled sheets
- Interpretation of Neural Networks is Fragile
- When will the mist clear? On the Interpretability of Machine Learning for Medical Applications: a survey
- Learning domain-agnostic visual representation for computational pathology using medically-irrelevant style transfer augmentation
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Do Explanations Reflect Decisions? A Machine-centric Strategy to Quantify the Performance of Explainability Algorithms
- Counterfactual State Explanations for Reinforcement Learning Agents via Generative Deep Learning
- Building and Interpreting Deep Similarity Models
- NBDT: Neural-Backed Decision Trees
- CEST MR fingerprinting (CEST-MRF) for Brain Tumor Quantification Using EPI Readout and Deep Learning Reconstruction
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Evaluating Explainable AI on a Multi-Modal Medical Imaging Task: Can Existing Algorithms Fulfill Clinical Requirements?
- Neuron Shapley: Discovering the Responsible Neurons
- Towards Relatable Explainable AI with the Perceptual Process
- On the (In)fidelity and Sensitivity for Explanations
- Machine Learning and value generation in Software Development: a survey
- Explainability Techniques for Graph Convolutional Networks
- A Survey of Deep Learning for Scientific Discovery
- Rationalization for Explainable NLP: A Survey
- Explanation-Guided Backdoor Poisoning Attacks Against Malware Classifiers
- Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
- Testing and verification of neural-network-based safety-critical control software: A systematic literature review
- Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps
- Real-time Neural Network Inference on Extremely Weak Devices: Agile Offloading with Explainable AI
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- A Spatio-Temporal Spot-Forecasting Framework for Urban Traffic Prediction
- Restricting the Flow: Information Bottlenecks for Attribution
- On Interpretability of Deep Learning based Skin Lesion Classifiers using Concept Activation Vectors
- On the Binding Problem in Artificial Neural Networks
- Beyond saliency: understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Temporal Graph Convolutional Networks for Automatic Seizure Detection
- Efficient Saliency Maps for Explainable AI
- Towards Evaluating Explanations of Vision Transformers for Medical Imaging
- Interpreting deep learning models for weak lensing
- On Interpretability of Artificial Neural Networks: A Survey
- Fast TreeSHAP: Accelerating SHAP Value Computation for Trees
- Convolutional Neural Nets in Chemical Engineering: Foundations, Computations, and Applications
- Symbolic Execution for Deep Neural Networks
- Grad-SAM: Explaining Transformers via Gradient Self-Attention Maps
- When Explanations Lie: Why Many Modified BP Attributions Fail
- Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
- Stakeholders in Explainable AI
- Getting a CLUE: A Method for Explaining Uncertainty Estimates
- Causality Learning: A New Perspective for Interpretable Machine Learning
- Amazon SageMaker Clarify: Machine Learning Bias Detection and Explainability in the Cloud
- MAGIX: Model Agnostic Globally Interpretable Explanations
- Adversarial Attacks and Defenses on Graphs: A Review, A Tool and Empirical Studies
- Understanding Neural Code Intelligence Through Program Simplification
- Regularizing Black-box Models for Improved Interpretability
- Towards Better Understanding Attribution Methods
- Invariant Rationalization
- Interpreting Graph Neural Networks for NLP With Differentiable Edge Masking
- NormLime: A New Feature Importance Metric for Explaining Deep Neural Networks
- Shedding Light on Black Box Machine Learning Algorithms: Development of an Axiomatic Framework to Assess the Quality of Methods that Explain Individual Predictions
- The many Shapley values for model explanation
- Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
- Photometric Redshifts from SDSS Images with an Interpretable Deep Capsule Network
- How does this interaction affect me? Interpretable attribution for feature interactions
- Global Extreme Heat Forecasting Using Neural Weather Models
- Rethinking Self-driving: Multi-task Knowledge for Better Generalization and Accident Explanation Ability
- Neural network interpretation using descrambler groups
- Group-CAM: Group Score-Weighted Visual Explanations for Deep Convolutional Networks
- Invertible Network for Classification and Biomarker Selection for ASD
- Neural Network Attributions: A Causal Perspective
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
- Explainable Machine Learning with Prior Knowledge: An Overview
- Understanding Integrated Gradients with SmoothTaylor for Deep Neural Network Attribution
- IS-CAM: Integrated Score-CAM for axiomatic-based explanations
- Code Prediction by Feeding Trees to Transformers
- A Survey on Neural Network Interpretability
- Towards falsifiable interpretability research
- On the Art and Science of Machine Learning Explanations
- White Box Methods for Explanations of Convolutional Neural Networks in Image Classification Tasks
- ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection
- Monitoring and explainability of models in production
- Shapley Flow: A Graph-based Approach to Interpreting Model Predictions
- Model-Based Counterfactual Synthesizer for Interpretation
- Towards Best Practice of Interpreting Deep Learning Models for EEG-based Brain Computer Interfaces
- SHAP values for Explaining CNN-based Text Classification Models
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- Weakly Supervised Deep Learning for COVID-19 Infection Detection and Classification from CT Images
- Counterfactual Explanation Based on Gradual Construction for Deep Networks
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- The effectiveness of feature attribution methods and its correlation with automatic evaluation scores
- Search for R-parity violating supersymmetry in a final state containing leptons and many jets with the ATLAS experiment using TeV proton-proton collision data
- Interpretability Beyond Classification Output: Semantic Bottleneck Networks
- Explaining Time Series Predictions with Dynamic Masks
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
- Bridging Adversarial Robustness and Gradient Interpretability
- Smoothed Geometry for Robust Attribution
- ferret: a Framework for Benchmarking Explainers on Transformers
- HYDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural Networks
- Interpretable deep learning for nuclear deformation in heavy ion collisions
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
- What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
- Measured and projected beam backgrounds in the Belle II experiment at the SuperKEKB collider
- Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses
- Towards Frequency-Based Explanation for Robust CNN
- Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
- The Portiloop: a deep learning-based open science tool for closed-loop brain stimulation
- Self-Attention Attribution: Interpreting Information Interactions Inside Transformer
- Explaining Deep Neural Networks
- Fast Hierarchical Games for Image Explanations
- Gifsplanation via Latent Shift: A Simple Autoencoder Approach to Counterfactual Generation for Chest X-rays
- Interpretable Directed Diversity: Leveraging Model Explanations for Iterative Crowd Ideation
- Evaluating Weakly Supervised Object Localization Methods Right
- Explaining Convolutional Neural Networks using Softmax Gradient Layer-wise Relevance Propagation
- Saliency Methods for Explaining Adversarial Attacks
- Token Prediction as Implicit Classification to Identify LLM-Generated Text
- Contextualizing Hate Speech Classifiers with Post-hoc Explanation
- Unwrapping The Black Box of Deep ReLU Networks: Interpretability, Diagnostics, and Simplification
- I-SPLIT: Deep Network Interpretability for Split Computing
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
- Explanation-Guided Training for Cross-Domain Few-Shot Classification
- Realised Volatility Forecasting: Machine Learning via Financial Word Embedding
- Feature Removal Is a Unifying Principle for Model Explanation Methods
- Deep Learning-based Type Identification of Volumetric MRI Sequences
- DISCOVER: 2-D Multiview Summarization of Optical Coherence Tomography Angiography for Automatic Diabetic Retinopathy Diagnosis
- CommonsenseVIS: Visualizing and Understanding Commonsense Reasoning Capabilities of Natural Language Models
- Learning Global Pairwise Interactions with Bayesian Neural Networks
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network
- Predicting Events in MOBA Games: Prediction, Attribution, and Evaluation
- Interpretable Deep Learning under Fire
- Deep Learning in Protein Structural Modeling and Design
- xGEMs: Generating Examplars to Explain Black-Box Models
- FairCanary: Rapid Continuous Explainable Fairness
- Refining Language Models with Compositional Explanations
- Interpreting Super-Resolution Networks with Local Attribution Maps
- Interpretable Neural Architecture Search via Bayesian Optimisation with Weisfeiler-Lehman Kernels
- Adversarial Infidelity Learning for Model Interpretation
- Towards Robust, Locally Linear Deep Networks
- Object-aware Contrastive Learning for Debiased Scene Representation
- Discovering and Explaining the Representation Bottleneck of DNNs
- Certifiably Robust Interpretation in Deep Learning
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural Networks
- Understanding Bias in Machine Learning
- Weakly-Supervised Action Localization and Action Recognition using Global-Local Attention of 3D CNN
- Generative causal explanations of black-box classifiers
- ML-LOO: Detecting Adversarial Examples with Feature Attribution
- Explaining Bayesian Neural Networks
- Are Visual Explanations Useful? A Case Study in Model-in-the-Loop Prediction
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- A Peg-in-hole Task Strategy for Holes in Concrete
- Explainability-based Backdoor Attacks Against Graph Neural Networks
- Best of both worlds: local and global explanations with human-understandable concepts
- Explaining a Deep Reinforcement Learning Docking Agent Using Linear Model Trees with User Adapted Visualization
- Fooling Partial Dependence via Data Poisoning
- Human Attention in Fine-grained Classification
- Exathlon: A Benchmark for Explainable Anomaly Detection over Time Series
- Blind Reverberation Time Estimation in Dynamic Acoustic Conditions
- Saliency is a Possible Red Herring When Diagnosing Poor Generalization
- Explaining Aggregates for Exploratory Analytics
- AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
- 3DB: A Framework for Debugging Computer Vision Models
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- EDUCE: Explaining model Decisions through Unsupervised Concepts Extraction
- On quantitative aspects of model interpretability
- Explaining a black-box using Deep Variational Information Bottleneck Approach
- Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation
- Comparative and Interpretative Analysis of CNN and Transformer Models in Predicting Wildfire Spread Using Remote Sensing Data
- The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- Benchmarking the Influence of Pre-training on Explanation Performance in MR Image Classification
- Same Words, Different Meanings: Semantic Polarization in Broadcast Media Language Forecasts Polarization on Social Media Discourse
- Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
- Analysis of Deep Networks for Monocular Depth Estimation Through Adversarial Attacks with Proposal of a Defense Method
- Joint Explainability and Sensitivity-Aware Federated Deep Learning for Transparent 6G RAN Slicing
- Why model why? Assessing the strengths and limitations of LIME
- Melody: Generating and Visualizing Machine Learning Model Summary to Understand Data and Classifiers Together
- On Network Science and Mutual Information for Explaining Deep Neural Networks
- AUTOLYCUS: Exploiting Explainable AI (XAI) for Model Extraction Attacks against Interpretable Models
- Interpretable Machine Learning: Moving From Mythos to Diagnostics
- Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes
- Deep learning the astrometric signature of dark matter substructure
- Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
- FUTURE-AI: Guiding Principles and Consensus Recommendations for Trustworthy Artificial Intelligence in Medical Imaging
- Revisiting Sanity Checks for Saliency Maps
- Towards Explainable Deep Learning for Credit Lending: A Case Study
- PredDiff: Explanations and Interactions from Conditional Expectations
- Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation
- Towards Understanding Neural Machine Translation with Word Importance
- Towards Interpretable Sparse Graph Representation Learning with Laplacian Pooling
- Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
- Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU Models
- Contrastive Explanations for Reinforcement Learning via Embedded Self Predictions
- MEGEX: Data-Free Model Extraction Attack against Gradient-Based Explainable AI
- Concise Explanations of Neural Networks using Adversarial Training
- Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis
- Fast Real-time Counterfactual Explanations
- Bias, Fairness, and Accountability with AI and ML Algorithms
- Sparse Oblique Decision Trees: A Tool to Understand and Manipulate Neural Net Features
- Nonet at SemEval-2023 Task 6: Methodologies for Legal Evaluation
- Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring
- Spatio-Temporal Perturbations for Video Attribution
- From Explanations to Segmentation: Using Explainable AI for Image Segmentation
- secml: A Python Library for Secure and Explainable Machine Learning
- What went wrong and when? Instance-wise Feature Importance for Time-series Models
- ECINN: Efficient Counterfactuals from Invertible Neural Networks
- A Note about: Local Explanation Methods for Deep Neural Networks lack Sensitivity to Parameter Values
- Explainable Deep Reinforcement Learning for UAV Autonomous Navigation
- Multi-modal Machine Learning for Vehicle Rating Predictions Using Image, Text, and Parametric Data
- Visualization of Convolutional Neural Networks for Monocular Depth Estimation
- Evaluating Input Perturbation Methods for Interpreting CNNs and Saliency Map Comparison
- ICAM: Interpretable Classification via Disentangled Representations and Feature Attribution Mapping
- Interpreting Deep Neural Networks Through Variable Importance
- Towards Aggregating Weighted Feature Attributions
- Learning to Identify Patients at Risk of Uncontrolled Hypertension Using Electronic Health Records Data
- IROF: a low resource evaluation metric for explanation methods
- Transformation Importance with Applications to Cosmology
- OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by learning to unfold
- Generalized Integrated Gradients: A practical method for explaining diverse ensembles
- Saliency-driven Word Alignment Interpretation for Neural Machine Translation
- T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
- Rationalizing Predictions by Adversarial Information Calibration
- CounteRGAN: Generating Realistic Counterfactuals with Residual Generative Adversarial Nets
- Computing a human-like reaction time metric from stable recurrent vision models
- Incorporating Priors with Feature Attribution on Text Classification
- i-Align: an interpretable knowledge graph alignment model
- This changes to that : Combining causal and non-causal explanations to generate disease progression in capsule endoscopy
- Improving the Transferability of Adversarial Attacks on Face Recognition with Beneficial Perturbation Feature Augmentation
- PointMask: Towards Interpretable and Bias-Resilient Point Cloud Processing
- Private Graph Extraction via Feature Explanations
- Does Your Model Think Like an Engineer? Explainable AI for Bearing Fault Detection with Deep Learning
- Algorithmic Recourse in the Wild: Understanding the Impact of Data and Model Shifts
- A Graph Based Neural Network Approach to Immune Profiling of Multiplexed Tissue Samples
- Interpretable and Accurate Fine-grained Recognition via Region Grouping
- RPGAN: GANs Interpretability via Random Routing
- What evidence does deep learning model use to classify Skin Lesions?
- Evaluation of Generalizability of Neural Program Analyzers under Semantic-Preserving Transformations
- DiaRet: A browser-based application for the grading of Diabetic Retinopathy with Integrated Gradients
- A Baseline for Shapley Values in MLPs: from Missingness to Neutrality
- On the Evaluation of the Plausibility and Faithfulness of Sentiment Analysis Explanations
- Provably efficient, succinct, and precise explanations
- Neutaint: Efficient Dynamic Taint Analysis with Neural Networks
- Evaluating neural network explanation methods using hybrid documents and morphological agreement
- Computationally Efficient Feature Significance and Importance for Machine Learning Models
- A simple defense against adversarial attacks on heatmap explanations
- On the Connection between Game-Theoretic Feature Attributions and Counterfactual Explanations
- Towards Interpretable Polyphonic Transcription with Invertible Neural Networks
- Human-interpretable model explainability on high-dimensional data
- Adversarial Attacks and Defenses: An Interpretation Perspective
- Why are Saliency Maps Noisy? Cause of and Solution to Noisy Saliency Maps
- Trusted Artificial Intelligence: Towards Certification of Machine Learning Applications
- EX-RAY: Distinguishing Injected Backdoor from Natural Features in Neural Networks by Examining Differential Feature Symmetry
- A Closer Look at Data Bias in Neural Extractive Summarization Models
- ProtoryNet - Interpretable Text Classification Via Prototype Trajectories
- Robustness of Visual Explanations to Common Data Augmentation
- Rethinking Natural Adversarial Examples for Classification Models
- CAUSE: Learning Granger Causality from Event Sequences using Attribution Methods
- Consistent Counterfactuals for Deep Models
- Interpreting and Boosting Dropout from a Game-Theoretic View
- The Struggles of Feature-Based Explanations: Shapley Values vs. Minimal Sufficient Subsets
- The Shapley Taylor Interaction Index
- XRAI: Better Attributions Through Regions
- Semantics for Global and Local Interpretation of Deep Neural Networks
- Technical Note: Game-Theoretic Interactions of Different Orders
- Deep Representations for Time-varying Brain Datasets
- Model Fusion via Optimal Transport
- Synthesizing Action Sequences for Modifying Model Decisions
- EvalAttAI: A Holistic Approach to Evaluating Attribution Maps in Robust and Non-Robust Models
- Debugging Tests for Model Explanations
- Order in the Court: Explainable AI Methods Prone to Disagreement
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations
- Hide-and-Seek: A Template for Explainable AI
- Cross-Modal Conceptualization in Bottleneck Models
- Robust Models Are More Interpretable Because Attributions Look Normal
- Usefulness of interpretability methods to explain deep learning based plant stress phenotyping
- Whatcha lookin' at? DeepLIFTing BERT's Attention in Question Answering
- SparCAssist: A Model Risk Assessment Assistant Based on Sparse Generated Counterfactuals
- Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence Models
- Interpretable Graph Capsule Networks for Object Recognition
- "How Does It Detect A Malicious App?" Explaining the Predictions of AI-based Android Malware Detector
- Evaluating Saliency Methods for Neural Language Models
- A Survey on the Robustness of Feature Importance and Counterfactual Explanations
- An Image Enhancing Pattern-based Sparsity for Real-time Inference on Mobile Devices
- Improving Feature Attribution through Input-specific Network Pruning
- Attribution Preservation in Network Compression for Reliable Network Interpretation
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Information-Theoretic Visual Explanation for Black-Box Classifiers
- Towards Interpretable Attention Networks for Cervical Cancer Analysis
- Computationally Efficient Measures of Internal Neuron Importance
- From Heatmaps to Structural Explanations of Image Classifiers
- One Explanation is Not Enough: Structured Attention Graphs for Image Classification
- Anomaly Attribution with Likelihood Compensation
- Maximally Invariant Data Perturbation as Explanation
- X-ToM: Explaining with Theory-of-Mind for Gaining Justified Human Trust
- Modelling EHR timeseries by restricting feature interaction
- Rethinking Cooperative Rationalization: Introspective Extraction and Complement Control
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations
- Distilling Ensemble of Explanations for Weakly-Supervised Pre-Training of Image Segmentation Models
- Generative Perturbation Analysis for Probabilistic Black-Box Anomaly Attribution
- Harmonizing Feature Attributions Across Deep Learning Architectures: Enhancing Interpretability and Consistency
- Massive s through the CNN lens: interpreting the field-level neutrino mass information in weak lensing
- Solving the enigma: Enhancing faithfulness and comprehensibility in explanations of deep networks
- On marginal feature attributions of tree-based models
- Optimising for Interpretability: Convolutional Dynamic Alignment Networks
- M2Lens: Visualizing and Explaining Multimodal Models for Sentiment Analysis
- Optimal MRI Undersampling Patterns for Ultimate Benefit of Medical Vision Tasks
- Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
- Can We Faithfully Represent Masked States to Compute Shapley Values on a DNN?
- Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity
- Robust and Stable Black Box Explanations
- Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
- Towards the Unification and Robustness of Perturbation and Gradient Based Explanations
- Improving Interpretability in Medical Imaging Diagnosis using Adversarial Training
- The Need for Standardized Explainability
- High Dimensional Model Explanations: an Axiomatic Approach
- Preserve, Promote, or Attack? GNN Explanation via Topology Perturbation
- Interpreting Interpretations: Organizing Attribution Methods by Criteria
- Interpretable Learning-to-Rank with Generalized Additive Models
- Learning Deep Attribution Priors Based On Prior Knowledge
- Rearchitecting Classification Frameworks For Increased Robustness
- Aggregating explanation methods for stable and robust explainability
- MonoNet: Towards Interpretable Models by Learning Monotonic Features
- Deep Neural Networks for Choice Analysis: Extracting Complete Economic Information for Interpretation
- A Simple Saliency Method That Passes the Sanity Checks
- Overcoming Catastrophic Forgetting by Generative Regularization
- A Review of Explainable Artificial Intelligence in Manufacturing
- Skillful Twelve Hour Precipitation Forecasts using Large Context Neural Networks
- TSInsight: A local-global attribution framework for interpretability in time-series data
- Towards Human-Interpretable Prototypes for Visual Assessment of Image Classification Models
- Discretized Integrated Gradients for Explaining Language Models
- Contextual Prediction Difference Analysis for Explaining Individual Image Classifications
- An Empirical Study of Explainable AI Techniques on Deep Learning Models For Time Series Tasks
- PhilaeX: Explaining the Failure and Success of AI Models in Malware Detection
- Towards Robust Explanations for Deep Neural Networks
- Deep Transfer Learning for Automated Diagnosis of Skin Lesions from Photographs
- Investigating Saturation Effects in Integrated Gradients
- Self-interpretable Convolutional Neural Networks for Text Classification
- Generative Counterfactuals for Neural Networks via Attribute-Informed Perturbation
- Gradient-based Analysis of NLP Models is Manipulable
- Evaluating and Characterizing Human Rationales
- Weakly Supervised Reasoning by Neuro-Symbolic Approaches
- When Fair Classification Meets Noisy Protected Attributes
- Unsupervised Representation Learning of DNA Sequences
- Visualizing Uncertainty and Saliency Maps of Deep Convolutional Neural Networks for Medical Imaging Applications
- Concept Learners for Few-Shot Learning
- Analyzing the Interpretability Robustness of Self-Explaining Models
- Unsupervised Detection of Distinctive Regions on 3D Shapes
- Deep Learning for Bug-Localization in Student Programs
- Feature Attributions and Counterfactual Explanations Can Be Manipulated
- Challenges for cognitive decoding using deep learning methods
- Visualizing Automatic Speech Recognition -- Means for a Better Understanding?
- Knowledge-based XAI through CBR: There is more to explanations than models can tell
- Rethinking Positive Aggregation and Propagation of Gradients in Gradient-based Saliency Methods
- This Reads Like That: Deep Learning for Interpretable Natural Language Processing
- WSAM: Visual Explanations from Style Augmentation as Adversarial Attacker and Their Influence in Image Classification
- Can Information Flows Suggest Targets for Interventions in Neural Circuits?
- On Robustness and Bias Analysis of BERT-based Relation Extraction
- Towards a Resilient Machine Learning Classifier -- a Case Study of Ransomware Detection
- The Penalty Imposed by Ablated Data Augmentation
- AES Systems Are Both Overstable And Oversensitive: Explaining Why And Proposing Defenses
- Gradient Weighted Superpixels for Interpretability in CNNs
- Assessment of the Reliablity of a Model's Decision by Generalizing Attribution to the Wavelet Domain
- Towards Rigorous Interpretations: a Formalisation of Feature Attribution
- Quantitative and Qualitative Evaluation of Explainable Deep Learning Methods for Ophthalmic Diagnosis
- Deriving Equivalent Symbol-Based Decision Models from Feedforward Neural Networks
- Conceptualizing Suicidal Behavior: Utilizing Explanations of Predicted Outcomes to Analyze Longitudinal Social Media Data
- SCOUT: Self-aware Discriminant Counterfactual Explanations
- Toward a Unified Framework for Debugging Concept-based Models
- Explaining Groups of Points in Low-Dimensional Representations
- A Causal Lens for Peeking into Black Box Predictive Models: Predictive Model Interpretation via Causal Attribution
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
- Computing Linear Restrictions of Neural Networks
- Born Identity Network: Multi-way Counterfactual Map Generation to Explain a Classifier's Decision
- On a Sparse Shortcut Topology of Artificial Neural Networks
- Attribution Analysis of Grammatical Dependencies in LSTMs
- DuTrust: A Sentiment Analysis Dataset for Trustworthiness Evaluation
- Quantitative Evaluation of Explainable Graph Neural Networks for Molecular Property Prediction
- FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging
- Learning from learning machines: a new generation of AI technology to meet the needs of science
- GANMEX: One-vs-One Attributions Guided by GAN-based Counterfactual Explanation Baselines
- Leveraging Sparse Linear Layers for Debuggable Deep Networks
- Towards Visually Explaining Video Understanding Networks with Perturbation
- XDeep: An Interpretation Tool for Deep Neural Networks
- Good-looking but Lacking Faithfulness: Understanding Local Explanation Methods through Trend-based Testing
- Uncertainty Propagation in Deep Neural Network Using Active Subspace
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- MATK: The Meme Analytical Tool Kit
- Algorithm Fairness in AI for Medicine and Healthcare
- CXR-Net: An Artificial Intelligence Pipeline for Quick Covid-19 Screening of Chest X-Rays
- Consensus-based Interpretable Deep Neural Networks with Application to Mortality Prediction
- What does LIME really see in images?
- Shapley Explanation Networks
- Explaining Image Classifiers using Statistical Fault Localization
- Scalable, Axiomatic Explanations of Deep Alzheimer's Diagnosis from Heterogeneous Data
- Deepfake Videos in the Wild: Analysis and Detection
- Switched linear projections for neural network interpretability
- Explaining Regression Based Neural Network Model
- Selective Ensembles for Consistent Predictions
- Toward the Understanding of Deep Text Matching Models for Information Retrieval
- Why Attentions May Not Be Interpretable?
- Energy-Based Learning for Cooperative Games, with Applications to Valuation Problems in Machine Learning
- Generative Visual Rationales
- Interpreting Dense Retrieval as Mixture of Topics
- Regularizing Black-box Models for Improved Interpretability (HILL 2019 Version)
- Evaluation of Saliency-based Explainability Method
- Adversarial TCAV -- Robust and Effective Interpretation of Intermediate Layers in Neural Networks
- Deep learning approaches for neural decoding: from CNNs to LSTMs and spikes to fMRI
- Measuring and improving the quality of visual explanations
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- A Comparison of Code Embeddings and Beyond
- Evaluation of importance estimators in deep learning classifiers for Computed Tomography
- Explaining neural network predictions of material strength
- iGOS++: Integrated Gradient Optimized Saliency by Bilateral Perturbations
- Towards Explanation of DNN-based Prediction with Guided Feature Inversion
- Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
- QUACKIE: A NLP Classification Task With Ground Truth Explanations
- Deep Learning for the Classification of Quenched Jets
- A Step Towards Exposing Bias in Trained Deep Convolutional Neural Network Models
- SkiNet: A Deep Learning Solution for Skin Lesion Diagnosis with Uncertainty Estimation and Explainability
- Model Interpretation and Explainability: Towards Creating Transparency in Prediction Models
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- Detecting cutaneous basal cell carcinomas in ultra-high resolution and weakly labelled histopathological images
- Investigating and Simplifying Masking-based Saliency Methods for Model Interpretability
- What Do Adversarially Robust Models Look At?
- Predicting sepsis in multi-site, multi-national intensive care cohorts using deep learning
- Studying Limits of Explainability by Integrated Gradients for Gene Expression Models
- FastEstimator: A Deep Learning Library for Fast Prototyping and Productization
- Understanding and Diagnosing Vulnerability under Adversarial Attacks
- Show or Suppress? Managing Input Uncertainty in Machine Learning Model Explanations
- VideoLightFormer: Lightweight Action Recognition using Transformers
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- A Unified Game-Theoretic Interpretation of Adversarial Robustness
- FastSHAP: Real-Time Shapley Value Estimation
- Partially Interpretable Estimators (PIE): Black-Box-Refined Interpretable Machine Learning
- IWA: Integrated Gradient based White-box Attacks for Fooling Deep Neural Networks
- Pre or Post-Softmax Scores in Gradient-based Attribution Methods, What is Best?
- SelfExplain: A Self-Explaining Architecture for Neural Text Classifiers
- On Sample Based Explanation Methods for NLP:Efficiency, Faithfulness, and Semantic Evaluation
- Unified Shapley Framework to Explain Prediction Drift
- Applicability Evaluation of Selected xAI Methods for Machine Learning Algorithms for Signal Parameters Extraction
- Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model Interpretation
- Exploring Self-Attention for Visual Odometry
- Where and When: Space-Time Attention for Audio-Visual Explanations
- Interpretable machine-learning identification of the crossover from subradiance to superradiance in an atomic array
- Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution
- Explanations for Occluded Images
- Do Input Gradients Highlight Discriminative Features?
- On the Robustness of Pretraining and Self-Supervision for a Deep Learning-based Analysis of Diabetic Retinopathy
- Who's a Good Boy? Reinforcing Canine Behavior in Real-Time using Machine Learning
- Robusta: Robust AutoML for Feature Selection via Reinforcement Learning
- Fairness-aware Summarization for Justified Decision-Making
- Geometry matters: Exploring language examples at the decision boundary
- BERTnesia: Investigating the capture and forgetting of knowledge in BERT
- The Eval4NLP Shared Task on Explainable Quality Estimation: Overview and Results
- Logic Traps in Evaluating Attribution Scores
- Exclusion and Inclusion -- A model agnostic approach to feature importance in DNNs
- Scaling Symbolic Methods using Gradients for Neural Model Explanation
- Making Document-Level Information Extraction Right for the Right Reasons
- Understanding Regularization to Visualize Convolutional Neural Networks
- Egocentric 6-DoF Tracking of Small Handheld Objects
- ExCon: Explanation-driven Supervised Contrastive Learning for Image Classification
- A Minimal Intervention Definition of Reverse Engineering a Neural Circuit
- Explainability-aided Domain Generalization for Image Classification
- Interpreting A Pre-trained Model Is A Key For Model Architecture Optimization: A Case Study On Wav2Vec 2.0
- Discriminative Attribution from Counterfactuals
- Natural Adversarial Objects
- DEPARA: Deep Attribution Graph for Deep Knowledge Transferability
- DANCE: Enhancing saliency maps using decoys
- Assessing the Reliability of Visual Explanations of Deep Models with Adversarial Perturbations
- Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
- Interpretable feature subset selection: A Shapley value based approach
- Gradient Frequency Modulation for Visually Explaining Video Understanding Models
- A Programmatic and Semantic Approach to Explaining and DebuggingNeural Network Based Object Detectors
- Uncertainty-aware Sensitivity Analysis Using Rényi Divergences
- Identifying Pediatric Vascular Anomalies With Deep Learning
- The Generalizability of Explanations
- Effective Use of Transformer Networks for Entity Tracking
- Physics-Inspired Interpretability Of Machine Learning Models
- Interpreting Undesirable Pixels for Image Classification on Black-Box Models
- Response Generation in Longitudinal Dialogues: Which Knowledge Representation Helps?
- Learning Invariances for Interpretability using Supervised VAE
- DRR4Covid: Learning Automated COVID-19 Infection Segmentation from Digitally Reconstructed Radiographs
- Pattern-Guided Integrated Gradients
- Human-in-the-loop model explanation via verbatim boundary identification in generated neighborhoods
- Machine Learning-Driven Analysis of kSZ Maps to Predict CMB Optical Depth
- On Spectral Properties of Gradient-based Explanation Methods
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Visualizing Color-wise Saliency of Black-Box Image Classification Models
- Interpretable Disentanglement of Neural Networks by Extracting Class-Specific Subnetwork
- Explaining Relation Classification Models with Semantic Extents
- Deep Q learning for fooling neural networks
- Online Black-Box Confidence Estimation of Deep Neural Networks
- On Attribution of Recurrent Neural Network Predictions via Additive Decomposition
- EDDA: Explanation-driven Data Augmentation to Improve Explanation Faithfulness
- A Peek Into the Reasoning of Neural Networks: Interpreting with Structural Visual Concepts
- A BERT-based Dual Embedding Model for Chinese Idiom Prediction
- SalKG: Learning From Knowledge Graph Explanations for Commonsense Reasoning
- A General Taylor Framework for Unifying and Revisiting Attribution Methods
- Longitudinal Distance: Towards Accountable Instance Attribution
- Explaining Away Attacks Against Neural Networks
- Saliency strikes back: How filtering out high frequencies improves white-box explanations
- Play Fair: Frame Attributions in Video Models
- A Novel Approach to Curiosity and Explainable Reinforcement Learning via Interpretable Sub-Goals
- Mutual Information Preserving Back-propagation: Learn to Invert for Faithful Attribution
- Modeling Cross-view Interaction Consistency for Paired Egocentric Interaction Recognition
- CNN-CASS: CNN for Classification of Coronary Artery Stenosis Score in MPR Images
- Corpus-level and Concept-based Explanations for Interpretable Document Classification
- A Case Study of Deep-Learned Activations via Hand-Crafted Audio Features
- Human-Imitating Metrics for Training and Evaluating Privacy Preserving Emotion Recognition Models Using Sociolinguistic Knowledge
- Interpreting Attributions and Interactions of Adversarial Attacks
- Emergent symbolic language based deep medical image classification
- Weakly Supervised Recovery of Semantic Attributes
- ProtoShotXAI: Using Prototypical Few-Shot Architecture for Explainable AI
- Noise Modulation: Let Your Model Interpret Itself
- Sanity Simulations for Saliency Methods
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- Learning Shape Features and Abstractions in 3D Convolutional Neural Networks for Detecting Alzheimer's Disease
- CACTUS: Detecting and Resolving Conflicts in Objective Functions
- Explainability-Aware One Point Attack for Point Cloud Neural Networks
- Human-Understandable Decision Making for Visual Recognition
- An exploration of the influence of path choice in game-theoretic attribution algorithms
- How to Explain Neural Networks: an Approximation Perspective
- Regularizing Explanations in Bayesian Convolutional Neural Networks
- Learning to Predict with Supporting Evidence: Applications to Clinical Risk Prediction
- Information-theoretic Evolution of Model Agnostic Global Explanations
- Interpretable Artificial Intelligence through the Lens of Feature Interaction
- Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning
- Explaining a prediction in some nonlinear models
- On Commonsense Cues in BERT for Solving Commonsense Tasks
- A Methodology for Exploring Deep Convolutional Features in Relation to Hand-Crafted Features with an Application to Music Audio Modeling
- Risk Prediction on Traffic Accidents using a Compact Neural Model for Multimodal Information Fusion over Urban Big Data
- SYSML: StYlometry with Structure and Multitask Learning: Implications for Darknet Forum Migrant Analysis
- Efficient Modelling Across Time of Human Actions and Interactions
- Attribution Mask: Filtering Out Irrelevant Features By Recursively Focusing Attention on Inputs of DNNs
- Reconstructing Actions To Explain Deep Reinforcement Learning
- Exploring Distantly-Labeled Rationales in Neural Network Models
- Interpretable Summaries of Black Box Incident Triaging with Subgroup Discovery
- Connecting Attributions and QA Model Behavior on Realistic Counterfactuals
- A Framework for Rationale Extraction for Deep QA models
- NeuroView: Explainable Deep Network Decision Making
- Robustness of different loss functions and their impact on networks learning capability
- Understanding Misclassifications by Attributes
- Inductive Granger Causal Modeling for Multivariate Time Series
- Unsupervised discovery of Interpretable Visual Concepts
- Discovering Invariances in Healthcare Neural Networks
- An Empirical Study towards Understanding How Deep Convolutional Nets Recognize Falls
- Causal Abstractions of Neural Networks
- Coalitional Bayesian Autoencoders -- Towards explainable unsupervised deep learning
- Multi-concept adversarial attacks
- Towards Explainable Fact Checking
- Separating Content and Style for Unsupervised Image-to-Image Translation
- Improving Attribution Methods by Learning Submodular Functions
- i-Algebra: Towards Interactive Interpretability of Deep Neural Networks
- Learning Propagation Rules for Attribution Map Generation
- Don't be fooled: label leakage in explanation methods and the importance of their quantitative evaluation
- diagNNose: A Library for Neural Activation Analysis
- Marginal Contribution Feature Importance -- an Axiomatic Approach for The Natural Case
- Evaluating Attribution Methods using White-Box LSTMs
- Semantic Network Interpretation
- Influence Patterns for Explaining Information Flow in BERT
- Signed Input Regularization
- Training Machine Learning Models by Regularizing their Explanations
- ILCRO: Making Importance Landscapes Flat Again
- Interpreting Deep Neural Networks with Relative Sectional Propagation by Analyzing Comparative Gradients and Hostile Activations
- Explainable Deep Modeling of Tabular Data using TableGraphNet
- Efficient Purely Convolutional Text Encoding
- Inserting Information Bottlenecks for Attribution in Transformers
- Towards Explainable Scientific Venue Recommendations
- Do not explain without context: addressing the blind spot of model explanations
- Automated Dependence Plots
- I know why you like this movie: Interpretable Efficient Multimodal Recommender
- Explainability Requires Interactivity
- Perturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
- How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs?
- Learning Robust Convolutional Neural Networks with Relevant Feature Focusing via Explanations
- To what extent do human explanations of model behavior align with actual model behavior?
- Explain and Improve: LRP-Inference Fine-Tuning for Image Captioning Models
- Explaining Classes through Word Attribution
- A Comparison of State-of-the-Art Techniques for Generating Adversarial Malware Binaries
- Games for Fairness and Interpretability
- Bayesian Interpolants as Explanations for Neural Inferences
- A Robust Unsupervised Ensemble of Feature-Based Explanations using Restricted Boltzmann Machines
- Enjoy the Salience: Towards Better Transformer-based Faithful Explanations with Word Salience
- Local Explanation of Dialogue Response Generation
- Thermostat: A Large Collection of NLP Model Explanations and Analysis Tools
- Explainable Deep Reinforcement Learning for Portfolio Management: An Empirical Approach
- It's FLAN time! Summing feature-wise latent representations for interpretability
- Exploring the Role of BERT Token Representations to Explain Sentence Probing Results
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog Systems
- DA-DGCEx: Ensuring Validity of Deep Guided Counterfactual Explanations With Distribution-Aware Autoencoder Loss
- Defense Against Explanation Manipulation
- Towards Auditability for Fairness in Deep Learning