14 citations · 20 across the 11 of their papers we have counts for
15 papers · 1 filter
Learning Beyond Limits: Multitask Learning and Synthetic Data for Low-Resource Canonical Morpheme Segmentation
Changbing Yang, Garrett Nicolai
We introduce a transformer-based morpheme segmentation system that augments a low-resource training signal through multitask learning and LLM-generated synthetic data. Our framewor…
Multiple Sources are Better Than One: Incorporating External Knowledge in Low-Resource Glossing
Changbing Yang, Garrett Nicolai, Miikka Silfverberg
In this paper, we address the data scarcity problem in automatic data-driven glossing for low-resource languages by coordinating multiple sources of linguistic expertise. We supple…
Embedded Translations for Low-resource Automated Glossing
Changbing Yang, Garrett Nicolai, Miikka Silfverberg
We investigate automatic interlinear glossing in low-resource settings. We augment a hard-attentional neural model with embedded translation information extracted from interlinear…
Neural Machine Translation Data Generation and Augmentation using ChatGPT
Wayne Yang, Garrett Nicolai
Neural models have revolutionized the field of machine translation, but creating parallel corpora is expensive and time-consuming. We investigate an alternative to manual parallel…
An Investigation of Noise in Morphological Inflection
Adam Wiemerslage, Changbing Yang, Garrett Nicolai +2
With a growing focus on morphological inflection systems for languages where high-quality data is scarce, training data noise is a serious but so far largely ignored concern. We ai…
Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Underdocumented Languages
Clarissa Forbes, Farhan Samir, Bruce Harold Oliver +4
Recent progress in NLP is driven by pretrained models leveraging massive datasets and has predominantly benefited the world's political and economic superpowers. Technologically un…