"Why Should I Trust You?": Explaining the Predictions of Any Classifier
arXiv:1602.04938
Abstract
Despite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one. In this work, we propose LIME, a novel explanation technique that explains the predictions of any classifier in an interpretable and faithful manner, by learning an interpretable model locally around the prediction. We also propose a method to explain models by presenting representative individual predictions and their explanations in a non-redundant way, framing the task as a submodular optimization problem. We demonstrate the flexibility of these methods by explaining different models for text (e.g. random forests) and image classification (e.g. neural networks). We show the utility of explanations via novel experiments, both simulated and with human subjects, on various scenarios that require trust: deciding if one should trust a prediction, choosing between models, improving an untrustworthy classifier, and identifying why a classifier should not be trusted.
References in corpus (1)
Cited by in corpus (77)
- Towards A Rigorous Science of Interpretable Machine Learning
- Axiomatic Attribution for Deep Networks
- Understanding Black-box Predictions via Influence Functions
- Machine Learning for Integrating Data in Biology and Medicine: Principles, Practice, and Opportunities
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Deep Convolutions for In-Depth Automated Rock Typing
- An unexpected unity among methods for interpreting model predictions
- Explainable AI: current status and future directions
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Text Classification using Capsules
- Machine Learning Model Interpretability for Precision Medicine
- Interpretation of Prediction Models Using the Input Gradient
- Data Science vs. Statistics: Two Cultures?
- Deep Learning for Computational Chemistry
- Three-stage intelligent support of clinical decision making for higher trust, validity, and explainability
- The Promise and Peril of Human Evaluation for Model Interpretability
- Locally Interpretable Models and Effects based on Supervised Partitioning (LIME-SUP)
- Interpretability of a Deep Learning Model in the Application of Cardiac MRI Segmentation with an ACDC Challenge Dataset
- The many Shapley values for model explanation
- Considerations When Learning Additive Explanations for Black-Box Models
- A Machine Learning Approach for Modelling Parking Duration in Urban Land-use
- Towards Safe Machine Learning for CPS: Infer Uncertainty from Training Data
- Can Explainable AI Explain Unfairness? A Framework for Evaluating Explainable AI
- Wearable Respiration Monitoring: Interpretable Inference with Context and Sensor Biomarkers
- "Influence Sketching": Finding Influential Samples In Large-Scale Regressions
- Knowledge Consistency between Neural Networks and Beyond
- Identifying Best Interventions through Online Importance Sampling
- Comparing Rule-Based and Deep Learning Models for Patient Phenotyping
- The Authority of "Fair" in Machine Learning
- Hollow-tree Super: a directional and scalable approach for feature importance in boosted tree models
- Representation Learning for Electronic Health Records
- Improving Simple Models with Confidence Profiles
- Poisoned classifiers are not only backdoored, they are fundamentally broken
- Quantitative Evaluations on Saliency Methods: An Experimental Study
- Zero-shot learning approach to adaptive Cybersecurity using Explainable AI
- Interpreting Embedding Models of Knowledge Bases: A Pedagogical Approach
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- Progressive Disclosure: Designing for Effective Transparency
- Impact of Accuracy on Model Interpretations
- Mapping chemical performance on molecular structures using locally interpretable explanations
- Closed-Form Expressions for Global and Local Interpretation of Tsetlin Machines with Applications to Explaining High-Dimensional Data
- Multiresolution Tensor Learning for Efficient and Interpretable Spatial Analysis
- Deep Learning in Information Security
- Principles of Explanation in Human-AI Systems
- XRAI: Better Attributions Through Regions
- Machine learning and behavioral economics for personalized choice architecture
- Fibres of Failure: Classifying errors in predictive processes
- Layer-wise Relevance Propagation for Explainable Recommendations
- Interpretable Learning-to-Rank with Generalized Additive Models
- Interpreting search result rankings through intent modeling
- Reliable Deep Grade Prediction with Uncertainty Estimation
- What Would You Ask the Machine Learning Model? Identification of User Needs for Model Explanations Based on Human-Model Conversations
- Enabling Machine Learning Algorithms for Credit Scoring -- Explainable Artificial Intelligence (XAI) methods for clear understanding complex predictive models
- Evaluating the performance of the LIME and Grad-CAM explanation methods on a LEGO multi-label image classification task
- Explanations of Machine Learning predictions: a mandatory step for its application to Operational Processes
- Explainable Recommender Systems via Resolving Learning Representations
- The Barrier of meaning in archaeological data science
- What's in the box? Explaining the black-box model through an evaluation of its interpretable features
- Class Introspection: A Novel Technique for Detecting Unlabeled Subclasses by Leveraging Classifier Explainability Methods
- Grounding Visual Explanations
- Machine Learning Algorithms for Financial Asset Price Forecasting
- How Interpretable and Trustworthy are GAMs?
- Interpretabilité des modèles : état des lieux des méthodes et application à l'assurance
- A Simple and Interpretable Predictive Model for Healthcare
- Towards Transparent Application of Machine Learning in Video Processing
- A Field Guide to Scientific XAI: Transparent and Interpretable Deep Learning for Bioinformatics Research
- Explanatory Pluralism in Explainable AI
- A Survey of Challenges and Opportunities in Sensing and Analytics for Cardiovascular Disorders
- "TL;DR:" Out-of-Context Adversarial Text Summarization and Hashtag Recommendation
- A Series of Unfortunate Counterfactual Events: the Role of Time in Counterfactual Explanations
- Axiomatic Interpretability for Multiclass Additive Models
- Self-learn to Explain Siamese Networks Robustly
- C2G-Net: Exploiting Morphological Properties for Image Classification
- Towards Interpretable Multi-Task Learning Using Bilevel Programming
- Theory In, Theory Out: The uses of social theory in machine learning for social science