5 papers
TREX: Tokenizer Regression for Optimal Data Mixture
Inho Won, Hangyeol Yoo, Minkyung Cho +3
Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer's compression performan…
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
Jayoung Song, KyungTae Lim, Jungyeul Park
Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the…
Enhancing Korean Dependency Parsing with Morphosyntactic Features
Jungyeul Park, Yige Chen, Kyuwon Kim +2
This paper introduces UniDive for Korean, an integrated framework that bridges Universal Dependencies (UD) and Universal Morphology (UniMorph) to enhance the representation and pro…
Parsing Through Boundaries in Chinese Word Segmentation
Yige Chen, Zelong Li, Cindy Zhang +7
Chinese word segmentation is a foundational task in natural language processing (NLP), with far-reaching effects on syntactic analysis. Unlike alphabetic languages like English, Ch…
K-UD: Revising Korean Universal Dependencies Guidelines
Kyuwon Kim, Yige Chen, Eunkyul Leah Jo +3
Critique has surfaced concerning the existing linguistic annotation framework for Korean Universal Dependencies (UDs), particularly in relation to syntactic relationships. In this…