Open Korean Corpora: A Practical Report
arXiv:2012.15621 · doi:10.18653/v1/2020.nlposs-1.12
Abstract
Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a list of Korean corpora, first describing institution-level resource development, then further iterate through a list of current open datasets for different types of tasks. We then propose a direction on how open-source dataset construction and releases should be done for less-resourced languages to promote research.
Published (v1) in NLP-OSS @EMNLP2020; May 2023 (v2) added with new datasets; June 2026 (v3) added analyses
References in corpus (18)
- Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Personalizing Dialogue Agents: I have a dog, do you have pets too?
- MultiCoNER: A Large-scale Multilingual dataset for Complex Named Entity Recognition
- KLUE: Korean Language Understanding Evaluation
- KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension
- KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding
- A Multi-Task Benchmark for Korean Legal Language Understanding and Judgement Prediction
- K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment
- Open Korean Corpora: A Practical Report
- Transformer-based Korean Pretrained Language Models: A Survey on Three Years of Progress
- Korean Online Hate Speech Dataset for Multilabel Classification: How Can Social Science Improve Dataset on Hate Speech?
- ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
- Pansori: ASR Corpus Generation from Open Online Video Contents
- StyleKQC: A Style-Variant Paraphrase Corpus for Korean Questions and Commands
- KOBEST: Korean Balanced Evaluation of Significant Tasks
- Speech Intention Understanding in a Head-final Language: A Disambiguation Utilizing Intonation-dependency
- KoCHET: a Korean Cultural Heritage corpus for Entity-related Tasks