3 papers
cs.CL2025
Linguistically Informed Tokenization Improves ASR for Underresourced Languages
Massimo Daul, Alessio Tosolini, Claire Bowern
Automatic speech recognition (ASR) is a crucial tool for linguists aiming to perform a variety of language documentation tasks. However, modern ASR systems use data-hungry transfor…
cs.CL2025
Multilingual MFA: Forced Alignment on Low-Resource Related Languages
Alessio Tosolini, Claire Bowern
We compare the outcomes of multilingual and crosslingual training for related and unrelated Australian languages with similar phonological inventories. We use the Montreal Forced A…
cs.CL2025
Data Augmentation and Hyperparameter Tuning for Low-Resource MFA
Alessio Tosolini, Claire Bowern
A continued issue for those working with computational tools and endangered and under-resourced languages is the lower accuracy of results for languages with smaller amounts of dat…