End-to-End QA on COVID-19: Domain Adaptation with Synthetic Training
arXiv:2012.01414
Abstract
End-to-end question answering (QA) requires both information retrieval (IR) over a large document collection and machine reading comprehension (MRC) on the retrieved passages. Recent work has successfully trained neural IR systems using only supervised question answering (QA) examples from open-domain datasets. However, despite impressive performance on Wikipedia, neural IR lags behind traditional term matching approaches such as BM25 in more specific and specialized target domains such as COVID-19. Furthermore, given little or no labeled data, effective adaptation of QA systems can also be challenging in such target domains. In this work, we explore the application of synthetically generated QA examples to improve performance on closed-domain retrieval and MRC. We combine our neural IR and MRC systems and show significant improvements in end-to-end QA on the CORD-19 collection over a state-of-the-art open-domain QA baseline.
Preprint
References in corpus (8)
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- CORD-19: The COVID-19 Open Research Dataset
- REALM: Retrieval-Augmented Language Model Pre-Training
- Rapidly Bootstrapping a Question Answering Dataset for COVID-19
- RikiNet: Reading Wikipedia Pages for Natural Question Answering
- Synthetic QA Corpora Generation with Roundtrip Consistency
- Answering Questions on COVID-19 in Real-Time
Cited by in corpus (5)
- Biomedical Question Answering: A Survey of Approaches and Challenges
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- Open-Domain Question-Answering for COVID-19 and Other Emergent Domains
- COVIDRead: A Large-scale Question Answering Dataset on COVID-19
- Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval