Incorporating Commonsense Knowledge into Abstractive Dialogue Summarization via Heterogeneous Graph Networks
arXiv:2010.10044
Abstract
Abstractive dialogue summarization is the task of capturing the highlights of a dialogue and rewriting them into a concise version. In this paper, we present a novel multi-speaker dialogue summarizer to demonstrate how large-scale commonsense knowledge can facilitate dialogue understanding and summary generation. In detail, we consider utterance and commonsense knowledge as two different types of data and design a Dialogue Heterogeneous Graph Network (D-HGN) for modeling both information. Meanwhile, we also add speakers as heterogeneous nodes to facilitate information flow. Experimental results on the SAMSum dataset show that our model can outperform various methods. We also conduct zero-shot setting experiments on the Argumentative Dialogue Summary Corpus, the results show that our model can better generalized to the new domain.
References in corpus (5)
- Semi-Supervised Classification with Graph Convolutional Networks
- SummaRuNNer: A Recurrent Neural Network based Sequence Model for Extractive Summarization of Documents
- A Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pretraining
- Topic-aware Pointer-Generator Networks for Summarizing Spoken Conversations
- Masking Orchestration: Multi-task Pretraining for Multi-role Dialogue Representation Learning
Cited by in corpus (7)
- Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization
- Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs
- GL-GIN: Fast and Accurate Non-Autoregressive Model for Joint Multiple Intent Detection and Slot Filling
- Controllable Abstractive Dialogue Summarization with Sketch Supervision
- Enhancing Semantic Understanding with Self-supervised Methods for Abstractive Dialogue Summarization
- Dialogue Summarization with Supporting Utterance Flow Modeling and Fact Regularization
- Low-Resource Dialogue Summarization with Domain-Agnostic Multi-Source Pretraining