activity
20242026
collaborators

6 papers

cs.CL2026

Learning Constituent Headedness

Zeyao Qi, Yige Chen, KyungTae Lim +2

Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…

cs.CL2026

TREX: Tokenizer Regression for Optimal Data Mixture

Inho Won, Hangyeol Yoo, Minkyung Cho +3

Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer's compression performan…

cs.CL2025

Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring

Jayoung Song, KyungTae Lim, Jungyeul Park

Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the…

cs.CL2025

Enhancing Korean Dependency Parsing with Morphosyntactic Features

Jungyeul Park, Yige Chen, Kyuwon Kim +2

This paper introduces UniDive for Korean, an integrated framework that bridges Universal Dependencies (UD) and Universal Morphology (UniMorph) to enhance the representation and pro…

cs.CL2024

K-UD: Revising Korean Universal Dependencies Guidelines

Kyuwon Kim, Yige Chen, Eunkyul Leah Jo +3

Critique has surfaced concerning the existing linguistic annotation framework for Korean Universal Dependencies (UDs), particularly in relation to syntactic relationships. In this…

cs.CL2024

Unlocking Korean Verbs: A User-Friendly Exploration into the Verb Lexicon

Seohyun Song, Eunkyul Leah Jo, Yige Chen +7

The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore l…