KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension
arXiv:1909.07005
Abstract
Machine Reading Comprehension (MRC) is a task that requires machine to understand natural language and answer questions by reading a document. It is the core of automatic response technology such as chatbots and automatized customer supporting systems. We present Korean Question Answering Dataset(KorQuAD), a large-scale Korean dataset for extractive machine reading comprehension task. It consists of 70,000+ human generated question-answer pairs on Korean Wikipedia articles. We release KorQuAD1.0 and launch a challenge at https://KorQuAD.github.io to encourage the development of multilingual natural language processing research.
References in corpus (1)
Cited by in corpus (17)
- QA Dataset Explosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension
- KLUE: Korean Language Understanding Evaluation
- TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
- FQuAD: French Question Answering Dataset
- An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks
- KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding
- Open Korean Corpora: A Practical Report
- A Vietnamese Dataset for Evaluating Machine Reading Comprehension
- ParsiNLU: A Suite of Language Understanding Challenges for Persian
- What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers
- A Survey on Awesome Korean NLP Datasets
- Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
- Analyzing Zero-shot Cross-lingual Transfer in Supervised NLP Tasks
- Improving Cross-Lingual Reading Comprehension with Self-Training
- Sentence Extraction-Based Machine Reading Comprehension for Vietnamese
- GermanQuAD and GermanDPR: Improving Non-English Question Answering and Passage Retrieval
- Structured Pattern Pruning Using Regularization