MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents
arXiv:2109.12595 · doi:10.18653/v1/2021.emnlp-main.498
Abstract
We propose MultiDoc2Dial, a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents. Most previous works treat document-grounded dialogue modeling as a machine reading comprehension task based on a single given document or passage. In this work, we aim to address more realistic scenarios where a goal-oriented information-seeking conversation involves multiple topics, and hence is grounded on different documents. To facilitate such a task, we introduce a new dataset that contains dialogues grounded in multiple documents from four different domains. We also explore modeling the dialogue-based and document-based context in the dataset. We present strong baseline approaches and various experimental results, aiming to support further research efforts on such a task.
References in corpus (6)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- REALM: Retrieval-Augmented Language Model Pre-Training
- Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
- Distilling Knowledge from Reader to Retriever for Question Answering
- A Graph-guided Multi-round Retrieval Method for Conversational Open-domain Question Answering
- Improving Unsupervised Dialogue Topic Segmentation with Utterance-Pair Coherence Scoring