Towards A Rigorous Science of Interpretable Machine Learning
arXiv:1702.08608
Abstract
As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other criteria such as safety or non-discrimination. However, despite the interest in interpretability, there is very little consensus on what interpretable machine learning is and how it should be measured. In this position paper, we first define interpretability and describe when interpretability is needed (and when it is not). Next, we suggest a taxonomy for rigorous evaluation and expose open questions towards a more rigorous science of interpretable machine learning.
References in corpus (2)
Cited by in corpus (11)
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- SmoothGrad: removing noise by adding noise
- 'It's Reducing a Human Being to a Percentage'; Perceptions of Justice in Algorithmic Decisions
- Explainable AI: Beware of Inmates Running the Asylum Or: How I Learnt to Stop Worrying and Love the Social and Behavioural Sciences
- AI Safety Gridworlds
- A Formal Framework to Characterize Interpretability of Procedures
- "I know it when I see it". Visualization and Intuitive Interpretability
- Training Feedforward Neural Networks with Standard Logistic Activations is Feasible
- Formalizing Interruptible Algorithms for Human over-the-loop Analytics
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Maintaining The Humanity of Our Models