"What is Relevant in a Text Document?": An Interpretable Machine Learning Approach
arXiv:1612.07843 · doi:10.1371/journal.pone.0181142
Abstract
Text documents can be described by a number of abstract concepts such as semantic category, writing style, or sentiment. Machine learning (ML) models have been trained to automatically map documents to these abstract concepts, allowing to annotate very large text collections, more than could be processed by a human in a lifetime. Besides predicting the text's category very accurately, it is also highly desirable to understand how and why the categorization process takes place. In this paper, we demonstrate that such understanding can be achieved by tracing the classification decision back to individual words using layer-wise relevance propagation (LRP), a recently developed technique for explaining predictions of complex non-linear classifiers. We train two word-based ML models, a convolutional neural network (CNN) and a bag-of-words SVM classifier, on a topic categorization task and adapt the LRP method to decompose the predictions of these models onto words. Resulting scores indicate how much individual words contribute to the overall classification decision. This enables one to distill relevant information from text documents without an explicit semantic information extraction step. We further use the word-wise relevance scores for generating novel vector-based document representations which capture semantic information. Based on these document vectors, we introduce a measure of model explanatory power and show that, although the SVM and CNN models perform similarly in terms of classification accuracy, the latter exhibits a higher level of explainability which makes it more comprehensible for humans and potentially more useful for other applications.
19 pages, 7 figures
References in corpus (5)
Cited by in corpus (62)
- Methods for Interpreting and Understanding Deep Neural Networks
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
- Towards Explainable Artificial Intelligence
- Interpretability of machine learning based prediction models in healthcare
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
- Explaining the Unique Nature of Individual Gait Patterns with Deep Learning
- Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI
- From Clustering to Cluster Explanations via Neural Networks
- Towards Explaining Anomalies: A Deep Taylor Decomposition of One-Class Models
- Explainable Artificial Intelligence: a Systematic Review
- Counterfactual Explanation Algorithms for Behavioral and Textual Data
- Explaining and Interpreting LSTMs
- Automating the search for a patent's prior art with a full text similarity search
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- A study on the Interpretability of Neural Retrieval Models using DeepSHAP
- Explainable Artificial Intelligence for Bayesian Neural Networks: Towards trustworthy predictions of ocean dynamics
- Local Interpretations for Explainable Natural Language Processing: A Survey
- Explaining First Impressions: Modeling, Recognizing, and Explaining Apparent Personality from Videos
- Legal Document Classification: An Application to Law Area Prediction of Petitions to Public Prosecution Service
- SHAP values for Explaining CNN-based Text Classification Models
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Gaussian Process Regression with Local Explanation
- Explanation-Guided Training for Cross-Domain Few-Shot Classification
- Explaining Bayesian Neural Networks
- Feature construction using explanations of individual predictions
- Why model why? Assessing the strengths and limitations of LIME
- Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis
- Progressive Disclosure: Designing for Effective Transparency
- Model Explainability in Deep Learning Based Natural Language Processing
- Evaluating neural network explanation methods using hybrid documents and morphological agreement
- Exploring text datasets by visualizing relevant words
- On the Evaluation of the Plausibility and Faithfulness of Sentiment Analysis Explanations
- Why an Android App is Classified as Malware? Towards Malware Classification Interpretation
- Game-Theoretic Interpretability for Temporal Modeling
- Towards Robust Explanations for Deep Neural Networks
- Self-interpretable Convolutional Neural Networks for Text Classification
- Towards Interpretable Deep Learning Models for Knowledge Tracing
- Evaluating Explanation Methods for Neural Machine Translation
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Taming Nonconvexity in Kernel Feature Selection -- Favorable Properties of the Laplace Kernel
- Path Analysis for Effective Fault Localization in Deep Neural Networks
- Unveiling Black-boxes: Explainable Deep Learning Models for Patent Classification
- Attention Flows are Shapley Value Explanations
- A Comprehensive Review on Summarizing Financial News Using Deep Learning
- Ontology-based Interpretable Machine Learning for Textual Data
- How do Convolutional Neural Networks Learn Design?
- Simplifying the explanation of deep neural networks with sufficient and necessary feature-sets: case of text classification
- Human-grounded Evaluations of Explanation Methods for Text Classification
- Teaching Solid Mechanics to Artificial Intelligence: a fast solver for heterogeneous solids
- "I had a solid theory before but it's falling apart": Polarizing Effects of Algorithmic Transparency
- Improving Moderation of Online Discussions via Interpretable Neural Models
- Detection of Dataset Shifts in Learning-Enabled Cyber-Physical Systems using Variational Autoencoder for Regression
- Pruning Attention Heads of Transformer Models Using A* Search: A Novel Approach to Compress Big NLP Architectures
- Robustness Tests of NLP Machine Learning Models: Search and Semantically Replace
- Saliency Maps Generation for Automatic Text Summarization
- Explainable Prediction of Text Complexity: The Missing Preliminaries for Text Simplification
- Distilling neural networks into skipgram-level decision lists
- Evaluating Attribution Methods using White-Box LSTMs