6 papers
Learning Constituent Headedness
Zeyao Qi, Yige Chen, KyungTae Lim +2
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…
TREX: Tokenizer Regression for Optimal Data Mixture
Inho Won, Hangyeol Yoo, Minkyung Cho +3
Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer's compression performan…
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
Jayoung Song, KyungTae Lim, Jungyeul Park
Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the…
Enhancing Korean Dependency Parsing with Morphosyntactic Features
Jungyeul Park, Yige Chen, Kyuwon Kim +2
This paper introduces UniDive for Korean, an integrated framework that bridges Universal Dependencies (UD) and Universal Morphology (UniMorph) to enhance the representation and pro…
K-UD: Revising Korean Universal Dependencies Guidelines
Kyuwon Kim, Yige Chen, Eunkyul Leah Jo +3
Critique has surfaced concerning the existing linguistic annotation framework for Korean Universal Dependencies (UDs), particularly in relation to syntactic relationships. In this…
Unlocking Korean Verbs: A User-Friendly Exploration into the Verb Lexicon
Seohyun Song, Eunkyul Leah Jo, Yige Chen +7
The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore l…