MEME: Generating RNN Model Explanations via Model Extraction
arXiv:2012.06954
Abstract
Recurrent Neural Networks (RNNs) have achieved remarkable performance on a range of tasks. A key step to further empowering RNN-based approaches is improving their explainability and interpretability. In this work we present MEME: a model extraction approach capable of approximating RNNs with interpretable models represented by human-understandable concepts and their interactions. We demonstrate how MEME can be applied to two multivariate, continuous data case studies: Room Occupation Prediction, and In-Hospital Mortality Prediction. Using these case-studies, we show how our extracted models can be used to interpret RNNs both locally and globally, by approximating RNN decision-making via interpretable concept interactions.
Presented at the HAMLETS workshop at the 34th Conference on Neural Information Processing Systems (NeurIPS 2020)
References in corpus (4)
- Beyond Sparsity: Tree Regularization of Deep Models for Interpretability
- Learning Deterministic Weighted Automata with Queries and Counterexamples
- Discovering Subdimensional Motifs of Different Lengths in Large-Scale Multivariate Time Series
- Efficient Algorithms for Generating Provably Near-Optimal Cluster Descriptors for Explainability