activity
20242026
collaborators

8 papers

cs.CL2026

TELLME: Test-Enhanced Learning for Language Model Enrichment

Minjun Kim, Inho Won, Hyeonseok Lim +6

Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such…

cs.CL2026

Refining Word-Based Grammatical Error Annotation for L2 Korean

Jungyeul Park, Kyungtae Lim, Wonjun Oh +4

Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner errors. Postpositions and verb…

cs.CL2026

Learning Constituent Headedness

Zeyao Qi, Yige Chen, KyungTae Lim +2

Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…

cs.CL2026

TREX: Tokenizer Regression for Optimal Data Mixture

Inho Won, Hangyeol Yoo, Minkyung Cho +3

Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer's compression performan…

cs.CL2025

Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring

Jayoung Song, KyungTae Lim, Jungyeul Park

Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the…

cs.CL2025

Unlocking Korean Verbs: A User-Friendly Exploration into the Verb Lexicon

Seohyun Song, Eunkyul Leah Jo, Yige Chen +7

The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore l…