Local Interpretations for Explainable Natural Language Processing: A Survey
arXiv:2103.11072 · doi:10.1145/3649450
Abstract
As the use of deep learning techniques has grown across various fields over the past decade, complaints about the opaqueness of the black-box models have increased, resulting in an increased focus on transparency in deep learning models. This work investigates various methods to improve the interpretability of deep neural networks for Natural Language Processing (NLP) tasks, including machine translation and sentiment analysis. We provide a comprehensive discussion on the definition of the term interpretability and its various aspects at the beginning of this work. The methods collected and summarised in this survey are only associated with local interpretation and are specifically divided into three categories: 1) interpreting the model's predictions through related input features; 2) interpreting through natural language explanation; 3) probing the hidden states of models and word representations.
Accepted by ACM Computing Surveys
References in corpus (8)
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex Models
- Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment
- Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond
- SceneGATE: Scene-Graph based co-Attention networks for TExt visual question answering
- Are Training Resources Insufficient? Predict First Then Explain!
- Generating Rationales in Visual Question Answering
- Learning to Rationalize for Nonmonotonic Reasoning with Distant Supervision