Evaluating neural network explanation methods using hybrid documents and morphological agreement
arXiv:1801.06422
Abstract
The behavior of deep neural networks (DNNs) is hard to understand. This makes it necessary to explore post hoc explanation methods. We conduct the first comprehensive evaluation of explanation methods for NLP. To this end, we design two novel evaluation paradigms that cover two important classes of NLP problems: small context and large context problems. Both paradigms require no manual annotation and are therefore broadly applicable. We also introduce LIMSSE, an explanation method inspired by LIME that is designed for NLP. We show empirically that LIMSSE, LRP and DeepLIFT are the most effective explanation methods and recommend them for explaining DNNs in NLP.
References in corpus (10)
- Axiomatic Attribution for Deep Networks
- Learning Important Features Through Propagating Activation Differences
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Understanding Neural Networks through Representation Erasure
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- Quasi-Recurrent Neural Networks
- Ask the GRU: Multi-Task Learning for Deep Text Recommendations
- Automatic Rule Extraction from Long Short Term Memory Networks
- Interpretation of Prediction Models Using the Input Gradient
- A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning