Towards A Rigorous Science of Interpretable Machine Learning
arXiv:1702.08608
Abstract
As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other criteria such as safety or non-discrimination. However, despite the interest in interpretability, there is very little consensus on what interpretable machine learning is and how it should be measured. In this position paper, we first define interpretability and describe when interpretability is needed (and when it is not). Next, we suggest a taxonomy for rigorous evaluation and expose open questions towards a more rigorous science of interpretable machine learning.
References in corpus (2)
Cited by in corpus (64)
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- SmoothGrad: removing noise by adding noise
- 'It's Reducing a Human Being to a Percentage'; Perceptions of Justice in Algorithmic Decisions
- Explainability Fact Sheets: A Framework for Systematic Assessment of Explainable Approaches
- Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI
- Explainable AI: Beware of Inmates Running the Asylum Or: How I Learnt to Stop Worrying and Love the Social and Behavioural Sciences
- AI Safety Gridworlds
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges
- Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based Systems
- "Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
- A Hierarchy of Limitations in Machine Learning
- Concise Fuzzy System Modeling Integrating Soft Subspace Clustering and Sparse Learning
- Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Values Approximation
- The Role of Cooperation in Responsible AI Development
- Counterfactual Visual Explanations
- Machine Learning Software Engineering in Practice: An Industrial Case Study
- Evaluating Saliency Map Explanations for Convolutional Neural Networks: A User Study
- The Price of Interpretability
- Interpretable deep learning for nuclear deformation in heavy ion collisions
- "The Human Body is a Black Box": Supporting Clinical Decision-Making with Deep Learning
- Global Explanations of Neural Networks: Mapping the Landscape of Predictions
- A Human-Grounded Evaluation of SHAP for Alert Processing
- Fairwashing: the risk of rationalization
- Machine learning and AI research for Patient Benefit: 20 Critical Questions on Transparency, Replicability, Ethics and Effectiveness
- A Formal Framework to Characterize Interpretability of Procedures
- ViCE: Visual Counterfactual Explanations for Machine Learning Models
- Neural Logic Reinforcement Learning
- Explaining Explanations to Society
- Personalized explanation in machine learning: A conceptualization
- Generating User-friendly Explanations for Loan Denials using GANs
- Disentangled Attribution Curves for Interpreting Random Forests and Boosted Trees
- Hybrid Predictive Model: When an Interpretable Model Collaborates with a Black-box Model
- On the Semantic Interpretability of Artificial Intelligence Models
- VINE: Visualizing Statistical Interactions in Black Box Models
- Modeling Heterogeneity in Mode-Switching Behavior Under a Mobility-on-Demand Transit System: An Interpretable Machine Learning Approach
- Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
- Unexplainability and Incomprehensibility of Artificial Intelligence
- Do Transformer Attention Heads Provide Transparency in Abstractive Summarization?
- GAN-based Generation and Automatic Selection of Explanations for Neural Networks
- A New Approach for Explainable Multiple Organ Annotation with Few Data
- Explainable Deep Relational Networks for Predicting Compound-Protein Affinities and Contacts
- "I know it when I see it". Visualization and Intuitive Interpretability
- Teaching Responsible Data Science: Charting New Pedagogical Territory
- Do ML Experts Discuss Explainability for AI Systems? A discussion case in the industry for a domain-specific solution
- Analysis Methods in Neural Language Processing: A Survey
- Nonlinear Semi-Parametric Models for Survival Analysis
- Understanding the Prediction Mechanism of Sentiments by XAI Visualization
- Training Feedforward Neural Networks with Standard Logistic Activations is Feasible
- Teaching AI to Explain its Decisions Using Embeddings and Multi-Task Learning
- NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language
- Interpretable Question Answering on Knowledge Bases and Text
- The Mass, Fake News, and Cognition Security
- Optimal Explanations of Linear Models
- Artificial Intelligence for Pediatric Ophthalmology
- Formalizing Interruptible Algorithms for Human over-the-loop Analytics
- Sampling Prediction-Matching Examples in Neural Networks: A Probabilistic Programming Approach
- From Receptive to Productive: Learning to Use Confusing Words through Automatically Selected Example Sentences
- Regression Concept Vectors for Bidirectional Explanations in Histopathology
- A general framework for scientifically inspired explanations in AI
- Maintaining The Humanity of Our Models
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Explainable Deep Modeling of Tabular Data using TableGraphNet
- Towards Interpretable Deep Extreme Multi-label Learning