activity
20242026
collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2026

Lexically conditioned realization ambiguity in Korean predicate morphology

Wonjun Oh, KyungTae Lim, Jungyeul Park

This paper examines Korean surface realization as distinct from morphological analysis. It asks whether a sequence of canonical morphemes and grammatical category labels uniquely d…

cs.CL2026

Representing and Parsing Korean Constituency Structure at Different Levels of Granularity

Jungyeul Park, KyungTae Lim, Zihao Huang +3

Korean constituency parsing raises a representational challenge because the terminal units of a phrase-structure tree do not straightforwardly correspond to simple surface words. K…

cs.CL2026

TELLME: Test-Enhanced Learning for Language Model Enrichment

Minjun Kim, Inho Won, Hyeonseok Lim +6

Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such…

cs.CL2026

Refining Word-Based Grammatical Error Annotation for L2 Korean

Jungyeul Park, Kyungtae Lim, Wonjun Oh +4

Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner errors. Postpositions and verb…

cs.CL2026

Learning Constituent Headedness

Zeyao Qi, Yige Chen, KyungTae Lim +2

Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…

cs.CL2026

TREX: Tokenizer Regression for Optimal Data Mixture

Inho Won, Hangyeol Yoo, Minkyung Cho +3

Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer's compression performan…