8 papers
TELLME: Test-Enhanced Learning for Language Model Enrichment
Minjun Kim, Inho Won, Hyeonseok Lim +6
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such…
Refining Word-Based Grammatical Error Annotation for L2 Korean
Jungyeul Park, Kyungtae Lim, Wonjun Oh +4
Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner errors. Postpositions and verb…
Learning Constituent Headedness
Zeyao Qi, Yige Chen, KyungTae Lim +2
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…
TREX: Tokenizer Regression for Optimal Data Mixture
Inho Won, Hangyeol Yoo, Minkyung Cho +3
Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer's compression performan…
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
Jayoung Song, KyungTae Lim, Jungyeul Park
Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the…
Unlocking Korean Verbs: A User-Friendly Exploration into the Verb Lexicon
Seohyun Song, Eunkyul Leah Jo, Yige Chen +7
The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore l…