activity
20242026
collaborators

12 papers

cs.CL2026

TELLME: Test-Enhanced Learning for Language Model Enrichment

Minjun Kim, Inho Won, Hyeonseok Lim +6

Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such…

cs.CL2026

Annotating Korean adnominal ending constructions in corpus data: Beyond relative-clause identification

Jungyeul Park, Chulwoo Park

The Korean adnominal ending \texttt{ETM} occurs in diverse noun-modifying constructions, including relative-clause-like modifiers, adjectival and copular forms, bound-noun construc…

cs.CL2026

Constituency Structure over Eojeol in Korean Treebanks

Jungyeul Park, Chulwoo Park

The design of Korean constituency treebanks raises a central representational question concerning the choice of terminal units. Although Korean words are morphologically complex, t…

cs.CL2026

Better heads do not guarantee better binarized constituency parsing

Zeyao Qi, Yige Chen, Eitan Klinger +2

We revisit punctuation-aware tree binarization for constituency parsing and ask whether dependency-induced headedness improves binary parser supervision. Although learned heads sub…

cs.CL2026

Chinese Word Boundary Recovery through Character Alignment Projection

Lusha Wang, Yuchen Li, Su Yuan +1

Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the word boundaries assumed by dow…

cs.CL2026

Learning Constituent Headedness

Zeyao Qi, Yige Chen, KyungTae Lim +2

Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…