3 citations · 3 across the 4 of their papers we have counts for
5 papers
Linguistically Informed Tokenization Improves ASR for Underresourced Languages
Massimo Daul, Alessio Tosolini, Claire Bowern
Automatic speech recognition (ASR) is a crucial tool for linguists aiming to perform a variety of language documentation tasks. However, modern ASR systems use data-hungry transfor…
Multilingual MFA: Forced Alignment on Low-Resource Related Languages
Alessio Tosolini, Claire Bowern
We compare the outcomes of multilingual and crosslingual training for related and unrelated Australian languages with similar phonological inventories. We use the Montreal Forced A…
Data Augmentation and Hyperparameter Tuning for Low-Resource MFA
Alessio Tosolini, Claire Bowern
A continued issue for those working with computational tools and endangered and under-resourced languages is the lower accuracy of results for languages with smaller amounts of dat…
Topic Modeling in the Voynich Manuscript
Rachel Sterneck, Annie Polish, Claire Bowern
This article presents the results of investigations using topic modeling of the Voynich Manuscript (Beinecke MS408). Topic modeling is a set of computational methods which are used…
Semantic Change and Semantic Stability: Variation is Key
Claire Bowern
I survey some recent approaches to studying change in the lexicon, particularly change in meaning across phylogenies. I briefly sketch an evolutionary approach to language change a…