SalKG: Learning From Knowledge Graph Explanations for Commonsense Reasoning
arXiv:2104.08793
Abstract
Augmenting pre-trained language models with knowledge graphs (KGs) has achieved success on various commonsense reasoning tasks. However, for a given task instance, the KG, or certain parts of the KG, may not be useful. Although KG-augmented models often use attention to focus on specific KG components, the KG is still always used, and the attention mechanism is never explicitly taught which KG components should be used. Meanwhile, saliency methods can measure how much a KG feature (e.g., graph, node, path) influences the model to make the correct prediction, thus explaining which KG features are useful. This paper explores how saliency explanations can be used to improve KG-augmented models' performance. First, we propose to create coarse (Is the KG useful?) and fine (Which nodes/paths in the KG are useful?) saliency explanations. Second, to motivate saliency-based supervision, we analyze oracle KG-augmented models which directly use saliency explanations as extra inputs for guiding their attention. Third, we propose SalKG, a framework for KG-augmented models to learn from coarse and/or fine saliency explanations. Given saliency explanations created from a task's training set, SalKG jointly trains the model to predict the explanations, then solve the task by attending to KG features highlighted by the predicted explanations. On three commonsense QA benchmarks (CSQA, OBQA, CODAH) and a range of KG-augmented models, we show that SalKG can yield considerable performance gains -- up to 2.76% absolute improvement on CSQA.
NeurIPS 2021
References in corpus (15)
- Attention is not Explanation
- Understanding Neural Networks through Representation Erasure
- Deep Learning: A Critical Appraisal
- Self-Attention Graph Pooling
- Language Models as Knowledge Bases?
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Extraction of Salient Sentences from Labelled Documents
- Commonsense Knowledge Mining from Pretrained Models
- Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question Answering
- Towards Interpretable Natural Language Understanding with Explanations as Latent Variables
- LIREx: Augmenting Language Inference with Relevant Explanation
- Connecting the Dots: A Knowledgeable Path Generator for Commonsense Question Answering
- Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language?
- Learning Contextualized Knowledge Structures for Commonsense Reasoning
- The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?