10 citations · 52 across the 47 of their papers we have counts for
7 papers · 2 filters
Mathematically Modeling the Lexicon Entropy of Emergent Language
Brendon Boldt, David Mortensen
We formulate a stochastic process, FiLex, as a mathematical model of lexicon entropy in deep learning-based emergent language systems. Defining a model mathematically allows it to…
Data-adaptive Transfer Learning for Translation: A Case Study in Haitian and Jamaican
Nathaniel R. Robinson, Cameron J. Hogan, Nancy Fulda +1
Multilingual transfer techniques often improve low-resource machine translation (MT). Many of these techniques are applied without considering data characteristics. We show in the…
ASR2K: Speech Recognition for Around 2000 Languages without Audio
Xinjian Li, Florian Metze, David R Mortensen +2
Most recent speech recognition models rely on large supervised datasets, which are unavailable for many low-resource languages. In this work, we present a speech recognition pipeli…
When Is TTS Augmentation Through a Pivot Language Useful?
Nathaniel Robinson, Perez Ogayo, Swetha Gangu +2
Developing Automatic Speech Recognition (ASR) for low-resource languages is a challenge due to the small amount of transcribed audio data. For many such languages, audio and text a…
Modeling Emergent Lexicon Formation with a Self-Reinforcing Stochastic Process
Brendon Boldt, David Mortensen
We introduce FiLex, a self-reinforcing stochastic process which models finite lexicons in emergent language experiments. The central property of FiLex is that it is a self-reinforc…
Learning the Ordering of Coordinate Compounds and Elaborate Expressions in Hmong, Lahu, and Chinese
Chenxuan Cui, Katherine J. Zhang, David R. Mortensen
Coordinate compounds (CCs) and elaborate expressions (EEs) are coordinate constructions common in languages of East and Southeast Asia. Mortensen (2006) claims that (1) the linear…