Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs
arXiv:1801.05453
Abstract
The driving force behind the recent success of LSTMs has been their ability to learn complex and non-linear relationships. Consequently, our inability to describe these relationships has led to LSTMs being characterized as black boxes. To this end, we introduce contextual decomposition (CD), an interpretation algorithm for analysing individual predictions made by standard LSTMs, without any changes to the underlying model. By decomposing the output of a LSTM, CD captures the contributions of combinations of words or variables to the final prediction of an LSTM. On the task of sentiment analysis with the Yelp and SST data sets, we show that CD is able to reliably identify words and phrases of contrasting sentiment, and how they are combined to yield the LSTM's final prediction. Using the phrase-level labels in SST, we also demonstrate that CD is able to successfully extract positive and negative negations from an LSTM, something which has not previously been done.
Oral presentation at ICLR 2018
References in corpus (5)
- Learning Important Features Through Propagating Activation Differences
- Understanding Neural Networks through Representation Erasure
- An unexpected unity among methods for interpreting model predictions
- Automatic Rule Extraction from Long Short Term Memory Networks
- On the State of the Art of Evaluation in Neural Language Models
Cited by in corpus (50)
- Interpretable machine learning: definitions, methods, and applications
- Pathologies of Neural Models Make Interpretations Difficult
- NeuralHydrology -- Interpreting LSTMs in Hydrology
- Exploring Interpretable LSTM Neural Networks over Multi-Variable Data
- Hierarchical interpretations for neural network predictions
- Learning Memory Access Patterns
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
- Interpreting Graph Neural Networks for NLP With Differentiable Edge Masking
- Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
- Detecting Statistical Interactions from Neural Network Weights
- Reverse engineering recurrent networks for sentiment classification reveals line attractor dynamics
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
- Self-Explaining Structures Improve NLP Models
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
- Self-Attention Attribution: Interpreting Information Interactions Inside Transformer
- Interpretable Deep Learning under Fire
- A Unified Approach to Interpreting and Boosting Adversarial Transferability
- Can I trust you more? Model-Agnostic Hierarchical Explanations
- Interpretable Anomaly Detection with DIFFI: Depth-based Isolation Forest Feature Importance
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- A Comparative Study of Rule Extraction for Recurrent Neural Networks
- Disentangled Attribution Curves for Interpreting Random Forests and Boosted Trees
- Transformation Importance with Applications to Cosmology
- Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text Classification
- Technical Note: Game-Theoretic Interactions of Different Orders
- Interpreting and Boosting Dropout from a Game-Theoretic View
- ProtoryNet - Interpretable Text Classification Via Prototype Trajectories
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Explainable Multivariate Time Series Classification: A Deep Neural Network Which Learns To Attend To Important Variables As Well As Informative Time Intervals
- Machine Learning for Multimodal Electronic Health Records-based Research: Challenges and Perspectives
- Can We Faithfully Represent Masked States to Compute Shapley Values on a DNN?
- Gradient-based Analysis of NLP Models is Manipulable
- Assessing Phrasal Representation and Composition in Transformers
- Distance and Equivalence between Finite State Machines and Recurrent Neural Networks: Computational results
- What do Deep Networks Like to Read?
- Defining and Quantifying the Emergence of Sparse Concepts in DNNs
- Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
- A Unified Game-Theoretic Interpretation of Adversarial Robustness
- Exclusion and Inclusion -- A model agnostic approach to feature importance in DNNs
- Towards Learning an Unbiased Classifier from Biased Data via Conditional Adversarial Debiasing
- Logic Traps in Evaluating Attribution Scores
- Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
- Local Explanation of Dialogue Response Generation
- Explain and Improve: LRP-Inference Fine-Tuning for Image Captioning Models
- LSTMs Compose (and Learn) Bottom-Up
- Evaluating Attribution Methods using White-Box LSTMs
- Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space
- Improving Moderation of Online Discussions via Interpretable Neural Models