3 papers
cs.CL2021
What shall we do with an hour of data? Speech recognition for the un- and under-served languages of Common Voice
Francis M. Tyers, Josh Meyer
This technical report describes the methods and results of a three-week sprint to produce deployable speech recognition models for 31 under-served languages of the Common Voice pro…
cs.CL2021
Few-Shot Keyword Spotting in Any Language
Mark Mazumder, Colby Banbury, Josh Meyer +2
We introduce a few-shot transfer learning method for keyword spotting in any language. Leveraging open speech corpora in nine languages, we automate the extraction of a large multi…
cs.CL2019
Common Voice: A Massively-Multilingual Speech Corpus
Rosana Ardila, Megan Branson, Kelly Davis +7
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic…