QA Dataset Explosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension
arXiv:2107.12708 · doi:10.1145/3560260
Abstract
Alongside huge volumes of research on deep learning models in NLP in the recent years, there has been also much work on benchmark datasets needed to track modeling progress. Question answering and reading comprehension have been particularly prolific in this regard, with over 80 new datasets appearing in the past two years. This study is the largest survey of the field to date. We provide an overview of the various formats and domains of the current resources, highlighting the current lacunae for future work. We further discuss the current classifications of "skills" that question answering/reading comprehension systems are supposed to acquire, and propose a new taxonomy. The supplementary materials survey the current multilingual resources and monolingual resources for languages other than English, and we discuss the implications of over-focusing on English. The study is aimed at both practitioners looking for pointers to the wealth of existing data, and at researchers working on new resources.
Published in ACM Comput. Surv (2022). This version differs from the final version in that section 7 ("Languages") is not in the main paper rather than the supplementary materials
References in corpus (31)
- Language Models are Few-Shot Learners
- Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
- Large-scale Simple Question Answering with Memory Networks
- Assessing BERT's Syntactic Abilities
- Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge Bases
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering
- Embracing data abundance: BookTest Dataset for Reading Comprehension
- Look before you Hop: Conversational Question Answering over Knowledge Graphs Using Judicious Context Expansion
- Automatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering
- A Survey on Neural Machine Reading Comprehension
- Question Answering is a Format; When is it Useful?
- Introducing MANtIS: a novel Multi-Domain Information Seeking Dialogues Dataset
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- When Do You Need Billions of Words of Pretraining Data?
- TableQA: a Large-Scale Chinese Text-to-SQL Dataset for Table-Aware SQL Generation
- MultiReQA: A Cross-Domain Evaluation for Retrieval Question Answering Models
- ORB: An Open Reading Benchmark for Comprehensive Evaluation of Machine Reading Comprehension
- Towards Question Format Independent Numerical Reasoning: A Set of Prerequisite Tasks
- LAReQA: Language-agnostic answer retrieval from a multilingual pool
- TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
- A Dataset and Baselines for Visual Question Answering on Art
- Perhaps PTLMs Should Go to School -- A Task to Assess Open Book and Closed Book QA
- Project PIAF: Building a Native French Question-Answering Dataset
- A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
- HEAD-QA: A Healthcare Dataset for Complex Reasoning
- ArchivalQA: A Large-scale Benchmark Dataset for Open Domain Question Answering over Historical News Collections
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
- ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion
- TIMEDIAL: Temporal Commonsense Reasoning in Dialog
Cited by in corpus (6)
- A Survey on Legal Question Answering Systems
- ShortcutLens: A Visual Analytics Approach for Exploring Shortcuts in Natural Language Understanding Dataset
- Dataset vs Reality: Understanding Model Performance from the Perspective of Information Need
- Identifying Experts in Question & Answer Portals: A Case Study on Data Science Competencies in Reddit
- ConveRT for FAQ Answering
- Investigating Post-pretraining Representation Alignment for Cross-Lingual Question Answering