How to Explain Neural Networks: an Approximation Perspective
arXiv:2105.07831
Abstract
The lack of interpretability has hindered the large-scale adoption of AI technologies. However, the fundamental idea of interpretability, as well as how to put it into practice, remains unclear. We provide notions of interpretability based on approximation theory in this study. We first implement this approximation interpretation on a specific model (fully connected neural network) and then propose to use MLP as a universal interpreter to explain arbitrary black-box models. Extensive experiments demonstrate the effectiveness of our approach.
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Understanding Neural Networks Through Deep Visualization
- SmoothGrad: removing noise by adding noise
- Distilling a Neural Network Into a Soft Decision Tree
- Deep Learning for Case-Based Reasoning through Prototypes: A Neural Network that Explains Its Predictions
- A Survey on Neural Network Interpretability
- Understanding the Decision Boundary of Deep Neural Networks: An Empirical Study