Interpretable Explanations of Black Boxes by Meaningful Perturbation
arXiv:1704.03296 · doi:10.1109/ICCV.2017.371
Abstract
As machine learning algorithms are increasingly applied to high impact yet high risk tasks, such as medical diagnosis or autonomous driving, it is critical that researchers can explain how such algorithms arrived at their predictions. In recent years, a number of image saliency methods have been developed to summarize where highly complex neural networks "look" in an image for evidence for their predictions. However, these techniques are limited by their heuristic nature and architectural constraints. In this paper, we make two main contributions: First, we propose a general framework for learning different kinds of explanations for any black box algorithm. Second, we specialise the framework to find the part of an image most responsible for a classifier decision. Unlike previous works, our method is model-agnostic and testable because it is grounded in explicit and interpretable image perturbations.
Final camera-ready paper published at ICCV 2017 (Supplementary materials: http://openaccess.thecvf.com/content_ICCV_2017/supplemental/Fong_Interpretable_Explanations_of_ICCV_2017_supplemental.pdf)
References in corpus (1)
Cited by in corpus (75)
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Towards Explainable Artificial Intelligence
- From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI
- A Comprehensive Taxonomy for Explainable Artificial Intelligence: A Systematic Survey of Surveys on Methods and Concepts
- XGNN: Towards Model-Level Explanations of Graph Neural Networks
- Explaining the Unique Nature of Individual Gait Patterns with Deep Learning
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- Automatic Sleep Staging of EEG Signals: Recent Development, Challenges, and Future Directions
- CheXplain: Enabling Physicians to Explore and UnderstandData-Driven, AI-Enabled Medical Imaging Analysis
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications
- Deep learning on fundus images detects glaucoma beyond the optic disc
- Impossibility Theorems for Feature Attribution
- ProtoPShare: Prototype Sharing for Interpretable Image Classification and Similarity Discovery
- Ensembles of Convolutional Neural Networks models for pediatric pneumonia diagnosis
- Testing and verification of neural-network-based safety-critical control software: A systematic literature review
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Deep Weakly-Supervised Learning Methods for Classification and Localization in Histology Images: A Survey
- Survey of XAI in digital pathology
- Leveraging Rationales to Improve Human Task Performance
- On the Eigenvalues of Global Covariance Pooling for Fine-grained Visual Recognition
- Towards Better Understanding Attribution Methods
- Interpreting Black Box Models via Hypothesis Testing
- A Survey on Neural Network Interpretability
- Towards Best Practice of Interpreting Deep Learning Models for EEG-based Brain Computer Interfaces
- Regional Multi-scale Approach for Visually Pleasing Explanations of Deep Neural Networks
- Explaining Black-box Models for Biomedical Text Classification
- CNN Attention Guidance for Improved Orthopedics Radiographic Fracture Classification
- SLISEMAP: Supervised dimensionality reduction through local explanations
- TrojanZoo: Towards Unified, Holistic, and Practical Evaluation of Neural Backdoors
- Object Detector Differences when using Synthetic and Real Training Data
- Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based Explanation
- CLEVR-X: A Visual Reasoning Dataset for Natural Language Explanations
- Feature Perturbation Augmentation for Reliable Evaluation of Importance Estimators in Neural Networks
- Towards Counterfactual and Contrastive Explainability and Transparency of DCNN Image Classifiers
- Deeply Explain CNN via Hierarchical Decomposition
- Occlusion Sensitivity Analysis with Augmentation Subspace Perturbation in Deep Feature Space
- TSGB: Target-Selective Gradient Backprop for Probing CNN Visual Saliency
- Neural Insights for Digital Marketing Content Design
- Sparse Oblique Decision Trees: A Tool to Understand and Manipulate Neural Net Features
- Spatio-Temporal Perturbations for Video Attribution
- Evaluating Input Perturbation Methods for Interpreting CNNs and Saliency Map Comparison
- LIMEcraft: Handcrafted superpixel selection and inspection for Visual eXplanations
- Studying How to Efficiently and Effectively Guide Models with Explanations
- Interpretable Machine Learning for Survival Analysis
- BSED: Baseline Shapley-Based Explainable Detector
- RouteExplainer: An Explanation Framework for Vehicle Routing Problem
- ADI: Adversarial Dominating Inputs in Vertical Federated Learning Systems
- ADVISE: ADaptive Feature Relevance and VISual Explanations for Convolutional Neural Networks
- From Heatmaps to Structural Explanations of Image Classifiers
- Time is Not Enough: Time-Frequency based Explanation for Time-Series Black-Box Models
- Evaluating Model Explanations without Ground Truth
- Better Understanding Differences in Attribution Methods via Systematic Evaluations
- Generating Post-hoc Explanations for Skip-gram-based Node Embeddings by Identifying Important Nodes with Bridgeness
- Neural Generators of Sparse Local Linear Models for Achieving both Accuracy and Interpretability
- Analyzing Effects of Mixed Sample Data Augmentation on Model Interpretability
- Understanding the Dependence of Perception Model Competency on Regions in an Image
- FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
- Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders
- Progress in deep Markov State Modeling: Coarse graining and experimental data restraints
- PCIM: Learning Pixel Attributions via Pixel-wise Channel Isolation Mixing in High Content Imaging
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding
- On Spectral Properties of Gradient-based Explanation Methods
- Leveraging Local Structure for Improving Model Explanations: An Information Propagation Approach
- Minimal Sufficient Views: A DNN model making predictions with more evidence has higher accuracy
- Ablation Path Saliency
- Explaining AI-based Decision Support Systems using Concept Localization Maps
- Towards the Characterization of Representations Learned via Capsule-based Network Architectures
- Class-Dependent Perturbation Effects in Evaluating Time Series Attributions
- Explainable machine learning classification of \textit{Chandra} X-ray sources: SHAP analysis of multi-wavelength features
- Interpretable Quantile Regression by Optimal Decision Trees
- Learn to Rank: Visual Attribution by Learning Importance Ranking
- LUDO: Low-Latency Understanding of Deformable Objects using Point Cloud Occupancy Functions