12 papers
TELLME: Test-Enhanced Learning for Language Model Enrichment
Minjun Kim, Inho Won, Hyeonseok Lim +6
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such…
Annotating Korean adnominal ending constructions in corpus data: Beyond relative-clause identification
Jungyeul Park, Chulwoo Park
The Korean adnominal ending \texttt{ETM} occurs in diverse noun-modifying constructions, including relative-clause-like modifiers, adjectival and copular forms, bound-noun construc…
Constituency Structure over Eojeol in Korean Treebanks
Jungyeul Park, Chulwoo Park
The design of Korean constituency treebanks raises a central representational question concerning the choice of terminal units. Although Korean words are morphologically complex, t…
Better heads do not guarantee better binarized constituency parsing
Zeyao Qi, Yige Chen, Eitan Klinger +2
We revisit punctuation-aware tree binarization for constituency parsing and ask whether dependency-induced headedness improves binary parser supervision. Although learned heads sub…
Chinese Word Boundary Recovery through Character Alignment Projection
Lusha Wang, Yuchen Li, Su Yuan +1
Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the word boundaries assumed by dow…
Learning Constituent Headedness
Zeyao Qi, Yige Chen, KyungTae Lim +2
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurall…