387 citations · 1.2k across the 47 of their papers we have counts for
7 papers · 2 filters
Joint Unsupervised and Supervised Training for Multilingual ASR
Junwen Bai, Bo Li, Yu Zhang +4
Self-supervised training has shown promising gains in pretraining models and facilitating the downstream finetuning for speech recognition, like multilingual ASR. Most existing met…
SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
Ankur Bapna, Yu-an Chung, Nan Wu +7
Unsupervised pre-training is now the predominant approach for both text and speech understanding. Self-attention models pre-trained on large amounts of unannotated data have been h…
Injecting Text in Self-Supervised Speech Pretraining
Zhehuai Chen, Yu Zhang, Andrew Rosenberg +3
Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretrainin…
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
William Chan, Daniel Park, Chris Lee +3
We present SpeechStew, a speech recognition model that is trained on a combination of various publicly available speech recognition datasets: AMI, Broadcast News, Common Voice, Lib…
Scaling End-to-End Models for Large-Scale Multilingual ASR
Bo Li, Ruoming Pang, Tara N. Sainath +7
Building ASR models across many languages is a challenging multi-task learning problem due to large variations and heavily unbalanced data. Existing work has shown positive transfe…
PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS
Ye Jia, Heiga Zen, Jonathan Shen +2
This paper introduces PnG BERT, a new encoder model for neural TTS. This model is augmented from the original BERT model, by taking both phoneme and grapheme representations of tex…