16 citations · 16 across the 3 of their papers we have counts for
4 papers
Quantifying Language Variation Acoustically with Few Resources
Martijn Bartelds, Martijn Wieling
Deep acoustic models represent linguistic information based on massive amounts of data. Unfortunately, for regional languages and dialects such resources are mostly not available.…
Automated speech tools for helping communities process restricted-access corpora for language revival efforts
Nay San, Martijn Bartelds, Tolúlopé Ògúnrèmí +6
Many archival recordings of speech from endangered languages remain unannotated and inaccessible to community members and language learning programs. One bottleneck is the time-int…
Adapting Monolingual Models: Data can be Scarce when Language Similarity is High
Wietse de Vries, Martijn Bartelds, Malvina Nissim +1
For many (minority) languages, the resources needed to train large models are not available. We investigate the performance of zero-shot transfer learning with as little data as po…
Leveraging pre-trained representations to improve access to untranscribed speech from endangered languages
Nay San, Martijn Bartelds, Mitchell Browne +9
Pre-trained speech representations like wav2vec 2.0 are a powerful tool for automatic speech recognition (ASR). Yet many endangered languages lack sufficient data for pre-training…