Attention is not Explanation
arXiv:1902.10186
Abstract
Attention mechanisms have seen wide adoption in neural NLP models. In addition to improving predictive performance, these are often touted as affording transparency: models equipped with attention provide a distribution over attended-to input units, and this is often presented (at least implicitly) as communicating the relative importance of inputs. However, it is unclear what relationship exists between attention weights and model outputs. In this work, we perform extensive experiments across a variety of NLP tasks that aim to assess the degree to which attention weights provide meaningful `explanations' for predictions. We find that they largely do not. For example, learned attention weights are frequently uncorrelated with gradient-based measures of feature importance, and one can identify very different attention distributions that nonetheless yield equivalent predictions. Our findings show that standard attention modules do not provide meaningful explanations and should not be treated as though they do. Code for all experiments is available at https://github.com/successar/AttentionExplanation.
Accepted as NAACL 2019 Long Paper
References in corpus (1)
Cited by in corpus (48)
- A Survey on the Explainability of Supervised Machine Learning
- A General Survey on Attention Mechanisms in Deep Learning
- A Survey on Aspect-Based Sentiment Classification
- On the Explainability of Natural Language Processing Deep Models
- WT5?! Training Text-to-Text Models to Explain their Predictions
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- The State of Human-centered NLP Technology for Fact-checking
- Explain and Predict, and then Predict Again
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Grad-SAM: Explaining Transformers via Gradient Self-Attention Maps
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
- Self-Explaining Structures Improve NLP Models
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Adversarial Infidelity Learning for Model Interpretation
- Normalized Attention Without Probability Cage
- Evolving Attention with Residual Convolutions
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Understanding Multi-Head Attention in Abstractive Summarization
- A multi-component framework for the analysis and design of explainable artificial intelligence
- On the Evaluation of the Plausibility and Faithfulness of Sentiment Analysis Explanations
- Understanding How Encoder-Decoder Architectures Attend
- Interpretation of NLP models through input marginalization
- A Study of the Plausibility of Attention between RNN Encoders in Natural Language Inference
- An attention model to analyse the risk of agitation and urinary tract infections in people with dementia
- Are Interpretations Fairly Evaluated? A Definition Driven Pipeline for Post-Hoc Interpretability
- Challenges for cognitive decoding using deep learning methods
- Weakly Supervised Reasoning by Neuro-Symbolic Approaches
- Attention Flows are Shapley Value Explanations
- Interpretable Question Answering on Knowledge Bases and Text
- VisBERT: Hidden-State Visualizations for Transformers
- Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
- Collaborative Graph Learning with Auxiliary Text for Temporal Event Prediction in Healthcare
- Attention improves concentration when learning node embeddings
- Does Attention Mechanism Possess the Feature of Human Reading? A Perspective of Sentiment Classification Task
- Spatial--spectral FFPNet: Attention-Based Pyramid Network for Segmentation and Classification of Remote Sensing Images
- Language Model Evaluation in Open-ended Text Generation
- Towards explainable message passing networks for predicting carbon dioxide adsorption in metal-organic frameworks
- Enriched Annotations for Tumor Attribute Classification from Pathology Reports with Limited Labeled Data
- Harnessing value from data science in business: ensuring explainability and fairness of solutions
- Transformer-F: A Transformer network with effective methods for learning universal sentence representation
- Explainability-aided Domain Generalization for Image Classification
- Multi-Domain Transformer-Based Counterfactual Augmentation for Earnings Call Analysis
- Attention or memory? Neurointerpretable agents in space and time
- Social Media Unrest Prediction during the {COVID}-19 Pandemic: Neural Implicit Motive Pattern Recognition as Psychometric Signs of Severe Crises
- A Framework for Rationale Extraction for Deep QA models
- Comparative Study of Language Models on Cross-Domain Data with Model Agnostic Explainability
- Games for Fairness and Interpretability
- The Brownian motion in the transformer model