Unsupervised Speech Recognition
arXiv:2105.11084
Abstract
Despite rapid progress in the recent past, current speech recognition systems still require labeled training data which limits this technology to a small fraction of the languages spoken around the globe. This paper describes wav2vec-U, short for wav2vec Unsupervised, a method to train speech recognition models without any labeled data. We leverage self-supervised speech representations to segment unlabeled audio and learn a mapping from these representations to phonemes via adversarial training. The right representations are key to the success of our method. Compared to the best previous unsupervised work, wav2vec-U reduces the phoneme error rate on the TIMIT benchmark from 26.1 to 11.3. On the larger English Librispeech benchmark, wav2vec-U achieves a word error rate of 5.9 on test-other, rivaling some of the best published systems trained on 960 hours of labeled data from only two years ago. We also experiment on nine other languages, including low-resource languages such as Kyrgyz, Swahili and Tatar.
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- Attention-Based Models for Speech Recognition
- Sequence Transduction with Recurrent Neural Networks
- Dual Learning for Machine Translation
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Semi-Supervised Speech Recognition via Local Prior Matching
- Differentiable Weighted Finite-State Transducers
Cited by in corpus (10)
- Unsupervised Automatic Speech Recognition: A Review
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings
- Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
- UniSpeech at scale: An Empirical Study of Pre-training Method on Large-Scale Speech Recognition Dataset
- Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
- NWT: Towards natural audio-to-video generation with representation learning
- Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?
- Towards an Efficient Voice Identification Using Wav2Vec2.0 and HuBERT Based on the Quran Reciters Dataset