Self-Training for End-to-End Speech Recognition
arXiv:1909.09116 · doi:10.1109/ICASSP40776.2020.9054295
Abstract
We revisit self-training in the context of end-to-end speech recognition. We demonstrate that training with pseudo-labels can substantially improve the accuracy of a baseline model. Key to our approach are a strong baseline acoustic and language model used to generate the pseudo-labels, filtering mechanisms tailored to common errors from sequence-to-sequence models, and a novel ensemble approach to increase pseudo-label diversity. Experiments on the LibriSpeech corpus show that with an ensemble of four models and label filtering, self-training yields a 33.9% relative improvement in WER compared with a baseline trained on 100 hours of labelled data in the noisy speech setting. In the clean speech setting, self-training recovers 59.3% of the gap between the baseline and an oracle model, which is at least 93.8% relatively higher than what previous approaches can achieve.
To be published in the 45th IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2020
References in corpus (9)
- Neural Machine Translation by Jointly Learning to Align and Translate
- Sequence to Sequence Learning with Neural Networks
- NLTK: The Natural Language Toolkit
- Cost-Effective Active Learning for Deep Image Classification
- Improved training of end-to-end attention models for speech recognition
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- wav2letter++: The Fastest Open-source Speech Recognition System
- RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation
- Fully Convolutional Speech Recognition
Cited by in corpus (41)
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- Rethinking Pre-training and Self-training
- Improved Noisy Student Training for Automatic Speech Recognition
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
- Self-Training: A Survey
- Effectiveness of self-supervised pre-training for speech recognition
- Margin Preserving Self-paced Contrastive Learning Towards Domain Adaptation for Medical Image Segmentation
- Iterative Pseudo-Labeling for Speech Recognition
- Unsupervised Speech Recognition
- On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR
- PseudoAugment: Learning to Use Unlabeled Data for Data Augmentation in Point Clouds
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
- MediaSpeech: Multilanguage ASR Benchmark and Dataset
- Learning Robust and Multilingual Speech Representations
- Self-training and Pre-training are Complementary for Speech Recognition
- Robust Mutual Learning for Semi-supervised Semantic Segmentation
- Magic dust for cross-lingual adaptation of monolingual wav2vec-2.0
- XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition
- Large-Scale Self- and Semi-Supervised Learning for Speech Translation
- Semi-supervised ASR by End-to-end Self-training
- Towards Semi-Supervised Semantics Understanding from Speech
- End-to-End Open Vocabulary Keyword Search With Multilingual Neural Representations
- Self-Training for End-to-End Speech Translation
- Momentum Pseudo-Labeling for Semi-Supervised Speech Recognition
- Data Transfer Approaches to Improve Seq-to-Seq Retrosynthesis
- Pseudo-labeling for Scalable 3D Object Detection
- Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
- Semi-Supervised Learning with Data Augmentation for End-to-End ASR
- Adversarial Meta Sampling for Multilingual Low-Resource Speech Recognition
- Contrastive Semi-supervised Learning for ASR
- Exploiting Large-scale Teacher-Student Training for On-device Acoustic Models
- ImageNet Pre-training also Transfers Non-Robustness
- AlphaMatch: Improving Consistency for Semi-supervised Learning with Alpha-divergence
- Unsupervised Adaptive Semantic Segmentation with Local Lipschitz Constraint
- Self-Training the Neurochaos Learning Algorithm
- Robust Generalization Strategies for Morpheme Glossing in an Endangered Language Documentation Context
- Multi-Object Tracking with Hallucinated and Unlabeled Videos
- Semi-Supervised Training with Pseudo-Labeling for End-to-End Neural Diarization