13 citations · 15 across the 7 of their papers we have counts for
7 papers
Pretrained self-supervised speech models can recognize unseen consonants
Chihiro Taguchi, Éric Le Ferrand, Hirosi Nakagawa +4
Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their tra…
Data-driven Parsing Evaluation for Child-Parent Interactions
Zoey Liu, Emily Prud'hommeaux
We present a syntactic dependency treebank for naturalistic child and child-directed speech in English (MacWhinney, 2000). Our annotations largely followed the guidelines of the Un…
Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation
Zoey Liu, Justin Spence, Emily Prud'hommeaux
Many automatic speech recognition (ASR) data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. This "hol…
Not always about you: Prioritizing community needs when developing endangered language technology
Zoey Liu, Crystal Richardson, Richard Hatcher +1
Languages are classified as low-resource when they lack the quantity of data necessary for training statistical and machine learning tools and models. Causes of resource scarcity v…
Data-driven Model Generalizability in Crosslinguistic Low-resource Morphological Segmentation
Zoey Liu, Emily Prud'hommeaux
Common designs of model evaluation typically focus on monolingual settings, where different models are compared according to their performance on a single data set that is assumed…
Neural Polysynthetic Language Modelling
Lane Schwartz, Francis Tyers, Lori Levin +18
Research in natural language processing commonly assumes that approaches that work well for English and and other widely-used languages are "language agnostic". In high-resource la…