activity
20182025
most citedThe SIGMORPHON 2019 Shared Task: Morphological Analysis in Context and Cross-Lingual Transfer for Inflection

14 citations · 20 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

15 papers · 1 filter

cs.CL2025

Learning Beyond Limits: Multitask Learning and Synthetic Data for Low-Resource Canonical Morpheme Segmentation

Changbing Yang, Garrett Nicolai

We introduce a transformer-based morpheme segmentation system that augments a low-resource training signal through multitask learning and LLM-generated synthetic data. Our framewor…

cs.CL2024

Multiple Sources are Better Than One: Incorporating External Knowledge in Low-Resource Glossing

Changbing Yang, Garrett Nicolai, Miikka Silfverberg

In this paper, we address the data scarcity problem in automatic data-driven glossing for low-resource languages by coordinating multiple sources of linguistic expertise. We supple…

cs.CL2024

Embedded Translations for Low-resource Automated Glossing

Changbing Yang, Garrett Nicolai, Miikka Silfverberg

We investigate automatic interlinear glossing in low-resource settings. We augment a hard-attentional neural model with embedded translation information extracted from interlinear…

cs.CL2023★ 2 cited

Neural Machine Translation Data Generation and Augmentation using ChatGPT

Wayne Yang, Garrett Nicolai

Neural models have revolutionized the field of machine translation, but creating parallel corpora is expensive and time-consuming. We investigate an alternative to manual parallel…

cs.CL2023★ 1 cited

An Investigation of Noise in Morphological Inflection

Adam Wiemerslage, Changbing Yang, Garrett Nicolai +2

With a growing focus on morphological inflection systems for languages where high-quality data is scarce, training data noise is a serious but so far largely ignored concern. We ai…

cs.CL2022

Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Underdocumented Languages

Clarissa Forbes, Farhan Samir, Bruce Harold Oliver +4

Recent progress in NLP is driven by pretrained models leveraging massive datasets and has predominantly benefited the world's political and economic superpowers. Technologically un…