Interpreting Deep Learning Models in Natural Language Processing: A Review
arXiv:2110.10470
Abstract
Neural network models have achieved state-of-the-art performances in a wide range of natural language processing (NLP) tasks. However, a long-standing criticism against neural network models is the lack of interpretability, which not only reduces the reliability of neural NLP systems but also limits the scope of their applications in areas where interpretability is essential (e.g., health care applications). In response, the increasing interest in interpreting neural NLP models has spurred a diverse array of interpretation methods over recent years. In this survey, we provide a comprehensive review of various interpretation methods for neural models in NLP. We first stretch out a high-level taxonomy for interpretation methods in NLP, i.e., training-based approaches, test-based approaches, and hybrid approaches. Next, we describe sub-categories in each category in detail, e.g., influence-function based methods, KNN-based methods, attention-based models, saliency-based methods, perturbation-based methods, etc. We point out deficiencies of current methods and suggest some avenues for future research.
References in corpus (28)
- Distilling the Knowledge in a Neural Network
- Language Models are Few-Shot Learners
- Understanding Neural Networks Through Deep Visualization
- A Structured Self-attentive Sentence Embedding
- Convolutional Neural Networks for Sentence Classification
- ERNIE: Enhanced Representation through Knowledge Integration
- SmoothGrad: removing noise by adding noise
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Attention is not Explanation
- Understanding Neural Networks through Representation Erasure
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Assessing BERT's Syntactic Abilities
- Towards a Human-like Open-Domain Chatbot
- Deep Learning for Case-Based Reasoning through Prototypes: A Neural Network that Explains Its Predictions
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Do Attention Heads in BERT Track Syntactic Dependencies?
- Extraction of Salient Sentences from Labelled Documents
- Attention Interpretability Across NLP Tasks
- A Game Theoretic Approach to Class-wise Selective Rationalization
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention
- The Price of Interpretability
- Self-Explaining Structures Improve NLP Models
- Unsupervised Latent Tree Induction with Deep Inside-Outside Recursive Autoencoders
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
- Inducing Syntactic Trees from BERT Representations
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Whatcha lookin' at? DeepLIFTing BERT's Attention in Question Answering