BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
arXiv:2109.13226 · doi:10.1109/JSTSP.2022.3182537
Abstract
We summarize the results of a host of efforts using giant automatic speech recognition (ASR) models pre-trained using large, diverse unlabeled datasets containing approximately a million hours of audio. We find that the combination of pre-training, self-training and scaling up model size greatly increases data efficiency, even for extremely large tasks with tens of thousands of hours of labeled data. In particular, on an ASR task with 34k hours of labeled data, by fine-tuning an 8 billion parameter pre-trained Conformer model we can match state-of-the-art (SoTA) performance with only 3% of the training data and significantly improve SoTA with the full training set. We also report on the universal benefits gained from using big pre-trained and self-trained models for a large set of downstream tasks that cover a wide range of speech domains and span multiple orders of magnitudes of dataset sizes, including obtaining SoTA performance on many public benchmarks. In addition, we utilize the learned representation of pre-trained networks to achieve SoTA results on non-ASR tasks.
14 pages, 7 figures, 13 tables; v2: minor corrections, reference baselines and bibliography updated; v3: corrections based on reviewer feedback, bibliography updated
References in corpus (22)
- Sequence Transduction with Recurrent Neural Networks
- On Using Monolingual Corpora in Neural Machine Translation
- Conformer: Convolution-augmented Transformer for Speech Recognition
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- An Unsupervised Autoregressive Model for Speech Representation Learning
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
- Multi-Format Contrastive Learning of Audio Representations
- Semi-Supervised Speech Recognition via Local Prior Matching
- Iterative Pseudo-Labeling for Speech Recognition
- On the importance of normative data in speech-based assessment
- Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training
- FRILL: A Non-Semantic Speech Embedding for Mobile Devices
- Scaling End-to-End Models for Large-Scale Multilingual ASR
- Learning Robust and Multilingual Speech Representations
- Spoken Language Identification using ConvNets
- Injecting Text in Self-Supervised Speech Pretraining
- Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
- Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models
Cited by in corpus (19)
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Measuring the Accuracy of Automatic Speech Recognition Solutions
- TRILLsson: Distilled Universal Paralinguistic Speech Representations
- Self-supervised representations in speech-based depression detection
- Universal Paralinguistic Speech Representations Using Self-Supervised Conformers
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition
- Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
- Parameter-Efficient Learning for Text-to-Speech Accent Adaptation
- How to Estimate Model Transferability of Pre-Trained Speech Models?
- Pre-Finetuning for Few-Shot Emotional Speech Recognition
- Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices
- Multilingual Speech Recognition using Knowledge Transfer across Learning Processes
- Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
- UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models