Towards Robust Interpretability with Self-Explaining Neural Networks
arXiv:1806.07538
Abstract
Most recent work on interpretability of complex machine learning models has focused on estimating explanations for previously trained models around specific predictions. models where interpretability plays a key role already during learning have received much less attention. We propose three desiderata for explanations in general -- explicitness, faithfulness, and stability -- and show that existing methods do not satisfy them. In response, we design self-explaining models in stages, progressively generalizing linear classifiers to complex yet architecturally explicit models. Faithfulness and stability are enforced via regularization specifically tailored to such models. Experimental results across various benchmark datasets show that our framework offers a promising direction for reconciling model complexity and interpretability.
NeurIPS 2018
References in corpus (4)
Cited by in corpus (118)
- A Survey on the Explainability of Supervised Machine Learning
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques
- Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations
- Adversarial attacks and defenses in explainable artificial intelligence: A survey
- Robust Machine Learning Systems: Challenges, Current Trends, Perspectives, and the Road Ahead
- Fooling Neural Network Interpretations via Adversarial Model Manipulation
- Acquisition of Chess Knowledge in AlphaZero
- TSViz: Demystification of Deep Learning Models for Time-Series Analysis
- How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations
- To trust or not to trust an explanation: using LEAF to evaluate local linear XAI methods
- Benchmarking Attribution Methods with Relative Feature Importance
- Explaining Anomalies Detected by Autoencoders Using SHAP
- Utilizing XAI technique to improve autoencoder based model for computer network anomaly detection with shapley additive explanation(SHAP)
- Counterfactual Explanations for Machine Learning on Multivariate Time Series Data
- Responsible and Regulatory Conform Machine Learning for Medicine: A Survey of Challenges and Solutions
- Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI
- Interpretable Deep Learning for Two-Prong Jet Classification with Jet Spectra
- On Interpretability of Artificial Neural Networks: A Survey
- Visualizing Deep Networks by Optimizing with Integrated Gradients
- A survey of algorithmic recourse: definitions, formulations, solutions, and prospects
- Regularizing Black-box Models for Improved Interpretability
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Neural Network Attributions: A Causal Perspective
- Explainability of Sub-Field Level Crop Yield Prediction using Remote Sensing
- Interpretable Deep Models for Cardiac Resynchronisation Therapy Response Prediction
- Stream-based active learning with linear models
- Interpretable Models for Granger Causality Using Self-explaining Neural Networks
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- Interpretability Beyond Classification Output: Semantic Bottleneck Networks
- Smoothed Geometry for Robust Attribution
- Self-Explaining Structures Improve NLP Models
- Born-Again Tree Ensembles
- Now You See Me (CME): Concept-based Model Extraction
- Embedding Deep Networks into Visual Explanations
- Gaussian Process Regression with Local Explanation
- Issues with post-hoc counterfactual explanations: a discussion
- U-Noise: Learnable Noise Masks for Interpretable Image Segmentation
- Fairwashing: the risk of rationalization
- Impossibility Results in AI: A Survey
- Certifiably Robust Interpretation in Deep Learning
- Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
- Generative causal explanations of black-box classifiers
- Explainable AI guided unsupervised fault diagnostics for high-voltage circuit breakers
- Interpretable Anomaly Detection with DIFFI: Depth-based Isolation Forest Feature Importance
- What Did You Think Would Happen? Explaining Agent Behaviour Through Intended Outcomes
- On quantitative aspects of model interpretability
- EDUCE: Explaining model Decisions through Unsupervised Concepts Extraction
- Landing AI on Networks: An equipment vendor viewpoint on Autonomous Driving Networks
- Understanding Instance-based Interpretability of Variational Auto-Encoders
- Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
- Rad4XCNN: a new agnostic method for post-hoc global explanation of CNN-derived features by means of radiomics
- CARLA: A Python Library to Benchmark Algorithmic Recourse and Counterfactual Explanation Algorithms
- LioNets: A Neural-Specific Local Interpretation Technique Exploiting Penultimate Layer Information
- timeXplain -- A Framework for Explaining the Predictions of Time Series Classifiers
- Regularizing Reasons for Outfit Evaluation with Gradient Penalty
- Interpretable and Accurate Fine-grained Recognition via Region Grouping
- Weight of Evidence as a Basis for Human-Oriented Explanations
- Human-interpretable model explainability on high-dimensional data
- XProtoNet: Diagnosis in Chest Radiography with Global and Local Explanations
- Benchmarks, Algorithms, and Metrics for Hierarchical Disentanglement
- Towards Robust Metrics for Concept Representation Evaluation
- Resisting Out-of-Distribution Data Problem in Perturbation of XAI
- MonoNet: Towards Interpretable Models by Learning Monotonic Features
- Weakly Supervised Multi-task Learning for Concept-based Explainability
- Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
- Optimising for Interpretability: Convolutional Dynamic Alignment Networks
- How Useful Are the Machine-Generated Interpretations to General Users? A Human Evaluation on Guessing the Incorrectly Predicted Labels
- DNN2LR: Interpretation-inspired Feature Crossing for Real-world Tabular Data
- Interpretable Text Classification Using CNN and Max-pooling
- An Empirical Study of Accuracy, Fairness, Explainability, Distributional Robustness, and Adversarial Robustness
- On Quantitative Evaluations of Counterfactuals
- Towards Interpretable Deep Learning Models for Knowledge Tracing
- TSInsight: A local-global attribution framework for interpretability in time-series data
- Neural Generators of Sparse Local Linear Models for Achieving both Accuracy and Interpretability
- Interpretable and Trustworthy Deepfake Detection via Dynamic Prototypes
- Concept Learners for Few-Shot Learning
- PatchX: Explaining Deep Models by Intelligible Pattern Patches for Time-series Classification
- Teaching the Machine to Explain Itself using Domain Knowledge
- Toward a Unified Framework for Debugging Concept-based Models
- A Quantitative Perspective on Values of Domain Knowledge for Machine Learning
- GANMEX: One-vs-One Attributions Guided by GAN-based Counterfactual Explanation Baselines
- Composition of Relational Features with an Application to Explaining Black-Box Predictors
- SelfExplain: A Self-Explaining Architecture for Neural Text Classifiers
- Learning interaction rules from multi-animal trajectories via augmented behavioral models
- Adherence and Constancy in LIME-RS Explanations for Recommendation
- Regularizing Black-box Models for Improved Interpretability (HILL 2019 Version)
- You Can Do Better! If You Elaborate the Reason When Making Prediction
- Shapley Explanation Networks
- A Semiparametric Approach to Interpretable Machine Learning
- Consensus-based Interpretable Deep Neural Networks with Application to Mortality Prediction
- Tree Boosted Varying Coefficient Models
- Interpretable by Design: Learning Predictors by Composing Interpretable Queries
- Learning Invariances for Interpretability using Supervised VAE
- Learning by Self-Explanation, with Application to Neural Architecture Search
- DANCE: Enhancing saliency maps using decoys
- The Definitions of Interpretability and Learning of Interpretable Models
- Towards Fully Interpretable Deep Neural Networks: Are We There Yet?
- Not All Features Are Equal: Feature Leveling Deep Neural Networks for Better Interpretation
- Explainability-aided Domain Generalization for Image Classification
- Bandits for Learning to Explain from Explanations
- The Partial Response Network: a neural network nomogram
- Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution
- Longitudinal Distance: Towards Accountable Instance Attribution
- Analyzing Representations inside Convolutional Neural Networks
- DNN2LR: Automatic Feature Crossing for Credit Scoring
- Learning to Predict with Supporting Evidence: Applications to Clinical Risk Prediction
- Learning Propagation Rules for Attribution Map Generation
- Weakly Supervised Recovery of Semantic Attributes
- Improving Attribution Methods by Learning Submodular Functions
- How to Explain Neural Networks: an Approximation Perspective
- Sanity Simulations for Saliency Methods
- Information-theoretic Evolution of Model Agnostic Global Explanations
- It's FLAN time! Summing feature-wise latent representations for interpretability
- Local Explanation of Dialogue Response Generation
- Deep Active Learning by Model Interpretability
- Defense Against Explanation Manipulation
- Self-explaining variational posterior distributions for Gaussian Process models