Attention Interpretability Across NLP Tasks
arXiv:1909.11218
Abstract
The attention layer in a neural network model provides insights into the model's reasoning behind its prediction, which are usually criticized for being opaque. Recently, seemingly contradictory viewpoints have emerged about the interpretability of attention weights (Jain & Wallace, 2019; Vig & Belinkov, 2019). Amid such confusion arises the need to understand attention mechanism more systematically. In this work, we attempt to fill this gap by giving a comprehensive explanation which justifies both kinds of observations (i.e., when is attention interpretable and when it is not). Through a series of experiments on diverse NLP tasks, we validate our observations and reinforce our claim of interpretability of attention through manual evaluation.
Cited by in corpus (9)
- On the Explainability of Natural Language Processing Deep Models
- Self-Explaining Structures Improve NLP Models
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Hard-Coded Gaussian Attention for Neural Machine Translation
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Understanding Multi-Head Attention in Abstractive Summarization
- A Simple Framework for Uncertainty in Contrastive Learning
- Enjoy the Salience: Towards Better Transformer-based Faithful Explanations with Word Salience
- Rationalizing Medical Relation Prediction from Corpus-level Statistics