Methods for Interpreting and Understanding Deep Neural Networks
arXiv:1706.07979 · doi:10.1016/j.dsp.2017.10.011
Abstract
This paper provides an entry point to the problem of interpreting a deep neural network model and explaining its predictions. It is based on a tutorial given at ICASSP 2017. It introduces some recently proposed techniques of interpretation, along with theory, tricks and recommendations, to make most efficient use of these techniques on real data. It also discusses a number of practical applications.
14 pages, 10 figures
References in corpus (4)
Cited by in corpus (116)
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
- A Survey on the Explainability of Supervised Machine Learning
- Machine learning for molecular simulation
- What Do We Want From Explainable Artificial Intelligence (XAI)? -- A Stakeholder Perspective on XAI and a Conceptual Model Guiding Interdisciplinary XAI Research
- Towards Explainable Artificial Intelligence
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Challenges of Real-World Reinforcement Learning
- A Review on Explainable Artificial Intelligence for Healthcare: Why, How, and When?
- Surveying the reach and maturity of machine learning and artificial intelligence in astronomy
- On the Explainability of Natural Language Processing Deep Models
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications
- Exploring Interpretable LSTM Neural Networks over Multi-Variable Data
- Explaining and Interpreting LSTMs
- Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment
- Concise Fuzzy System Modeling Integrating Soft Subspace Clustering and Sparse Learning
- Towards Explainable Neural-Symbolic Visual Reasoning
- Understanding and Comparing Deep Neural Networks for Age and Gender Classification
- Intrapapillary Capillary Loop Classification in Magnification Endoscopy: Open Dataset and Baseline Methodology
- Q-value Path Decomposition for Deep Multiagent Reinforcement Learning
- Explainable artificial intelligence model to predict acute critical illness from electronic health records
- Bias in Data-driven AI Systems -- An Introductory Survey
- Model extraction from counterfactual explanations
- Fairwashing: the risk of rationalization
- How Much Can I Trust You? -- Quantifying Uncertainties in Explaining Neural Networks
- Breaking Batch Normalization for better explainability of Deep Neural Networks through Layer-wise Relevance Propagation
- Domain Knowledge Aided Explainable Artificial Intelligence for Intrusion Detection and Response
- Calibrating Healthcare AI: Towards Reliable and Interpretable Deep Predictive Models
- A Rate-Distortion Framework for Explaining Neural Network Decisions
- Unbox the Black-box for the Medical Explainable AI via Multi-modal and Multi-centre Data Fusion: A Mini-Review, Two Showcases and Beyond
- Feature Visualization within an Automated Design Assessment leveraging Explainable Artificial Intelligence Methods
- Deep Representation Learning for Social Network Analysis
- The Convergence of Machine Learning and Communications
- Personalized explanation in machine learning: A conceptualization
- Revisiting Sanity Checks for Saliency Maps
- Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU Models
- Don't Explain without Verifying Veracity: An Evaluation of Explainable AI with Video Activity Recognition
- Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
- CDT: Cascading Decision Trees for Explainable Reinforcement Learning
- On Relating 'Why?' and 'Why Not?' Explanations
- Applying saliency-map analysis in searches for pulsars and fast radio bursts
- Model Explainability in Deep Learning Based Natural Language Processing
- Saliency-driven Word Alignment Interpretation for Neural Machine Translation
- Regularizing Reasons for Outfit Evaluation with Gradient Penalty
- Probing Criticality in Quantum Spin Chains with Neural Networks
- Exploring text datasets by visualizing relevant words
- Achievements and Challenges in Explaining Deep Learning based Computer-Aided Diagnosis Systems
- Debugging Tests for Model Explanations
- Efficient Search for Diverse Coherent Explanations
- Noise reduction on single-shot images using an autoencoder
- Complementary reinforcement learning towards explainable agents
- Causal Analysis of Agent Behavior for AI Safety
- Attribution Preservation in Network Compression for Reliable Network Interpretation
- DeepSeer: Interactive RNN Explanation and Debugging via State Abstraction
- Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence Models
- GAN-based Generation and Automatic Selection of Explanations for Neural Networks
- Harmonizing Feature Attributions Across Deep Learning Architectures: Enhancing Interpretability and Consistency
- A Simple Saliency Method That Passes the Sanity Checks
- From Shallow to Deep Interactions Between Knowledge Representation, Reasoning and Machine Learning (Kay R. Amel group)
- Heteroscedastic Calibration of Uncertainty Estimators in Deep Learning
- Preserve, Promote, or Attack? GNN Explanation via Topology Perturbation
- A Survey on Understanding, Visualizations, and Explanation of Deep Neural Networks
- Efficient Explanations With Relevant Sets
- The Skincare project, an interactive deep learning system for differential diagnosis of malignant skin lesions. Technical Report
- Towards Interpretable Deep Learning Models for Knowledge Tracing
- Evaluation, Tuning and Interpretation of Neural Networks for Meteorological Applications
- On Efficiently Explaining Graph-Based Classifiers
- Towards Robust Explanations for Deep Neural Networks
- Ensemble Transfer Learning of Elastography and B-mode Breast Ultrasound Images
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Scene Text Recognition Models Explainability Using Local Features
- Embedded Encoder-Decoder in Convolutional Networks Towards Explainable AI
- A Causal Lens for Peeking into Black Box Predictive Models: Predictive Model Interpretation via Causal Attribution
- Discovering topics in text datasets by visualizing relevant words
- EMAP: Explanation by Minimal Adversarial Perturbation
- Computer Vision and Deep Learning for Fish Classification in Underwater Habitats: A Survey
- Computing Optimal Decision Sets with SAT
- Teaching AI to Explain its Decisions Using Embeddings and Multi-Task Learning
- Location, location, location: Satellite image-based real-estate appraisal
- Explainable Recommender Systems via Resolving Learning Representations
- Interpretable Time-series Classification on Few-shot Samples
- Reachability Analysis of Convolutional Neural Networks
- Explaining the Road Not Taken
- Deep Learning Based Decision Support for Medicine -- A Case Study on Skin Cancer Diagnosis
- PLANS: Robust Program Learning from Neurally Inferred Specifications
- On the Robustness of Pretraining and Self-Supervision for a Deep Learning-based Analysis of Diabetic Retinopathy
- Inference of cell dynamics on perturbation data using adjoint sensitivity
- Interpretable Disentanglement of Neural Networks by Extracting Class-Specific Subnetwork
- An AI-Augmented Lesion Detection Framework For Liver Metastases With Model Interpretability
- Assessing the Reliability of Visual Explanations of Deep Models with Adversarial Perturbations
- On Explaining Random Forests with SAT
- Position Paper: Towards Transparent Machine Learning
- Explainability's Gain is Optimality's Loss? -- How Explanations Bias Decision-making
- Explaining Classes through Word Attribution
- An Interpretable Neural Network for Parameter Inference
- A crossover code for high-dimensional composition
- Deep Relevance Regularization: Interpretable and Robust Tumor Typing of Imaging Mass Spectrometry Data
- Coherence of Working Memory Study Between Deep Neural Network and Neurophysiology
- Regression Concept Vectors for Bidirectional Explanations in Histopathology
- Get It Scored Using AutoSAS -- An Automated System for Scoring Short Answers
- Distilling neural networks into skipgram-level decision lists
- Investigating ADR mechanisms with knowledge graph mining and explainable AI
- Convolutional Neural Networks from Image Markers
- Assessing The Importance Of Colours For CNNs In Object Recognition
- Data-Driven Continuum Dynamics via Transport-Teleport Duality
- Low-Dimensional Manifolds Support Multiplexed Integrations in Recurrent Neural Networks
- Focused LRP: Explainable AI for Face Morphing Attack Detection
- A Neural-Symbolic Framework for Mental Simulation
- ST-ABN: Visual Explanation Taking into Account Spatio-temporal Information for Video Recognition
- On Predictive Explanation of Data Anomalies
- Logic Constraints to Feature Importances
- Langevin Cooling for Domain Translation
- Neuro-Symbolic AI: An Emerging Class of AI Workloads and their Characterization
- Self-explaining variational posterior distributions for Gaussian Process models
- A User-Centred Framework for Explainable Artificial Intelligence in Human-Robot Interaction
- T3-Vis: a visual analytic framework for Training and fine-Tuning Transformers in NLP