5 papers
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo +3
We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to sp…
Investigating Transcription Normalization in the Faetar ASR Benchmark
Leo Peckham, Michael Ong, Naomi Nagy +1
We examine the role of transcription inconsistencies in the Faetar Automatic Speech Recognition benchmark, a challenging low-resource ASR benchmark. With the help of a small, hand-…
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
Robin Huo, Ewan Dunbar
Self-supervised models for speech representation learning now see widespread use for their versatility and performance on downstream tasks, but the effect of model architecture on…
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
Michael Ong, Sean Robertson, Leo Peckham +7
We introduce the Faetar Automatic Speech Recognition Benchmark, a benchmark corpus designed to push the limits of current approaches to low-resource speech recognition. Faetar, a F…
Quantifying the Role of Textual Predictability in Automatic Speech Recognition
Sean Robertson, Gerald Penn, Ewan Dunbar
A long-standing question in automatic speech recognition research is how to attribute errors to the ability of a model to model the acoustics, versus its ability to leverage higher…