859 citations · 1.1k across the 26 of their papers we have counts for
43 papers
Resource-Efficient Transfer Learning From Speech Foundation Model Using Hierarchical Feature Fusion
Zhouyuan Huo, Khe Chai Sim, Bo Li +3
Self-supervised pre-training of a speech foundation model, followed by supervised fine-tuning, has shown impressive quality improvements on automatic speech recognition (ASR) tasks…
JOIST: A Joint Speech and Text Streaming Model For ASR
Tara N. Sainath, Rohit Prabhavalkar, Ankur Bapna +6
We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous wo…
Scaling Up Deliberation for Multilingual ASR
Ke Hu, Bo Li, Tara N. Sainath
Multilingual end-to-end automatic speech recognition models are attractive due to its simplicity in training and deployment. Recent work on large-scale training of such models has…
Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification
Chao Zhang, Bo Li, Tara Sainath +4
Language identification is critical for many downstream tasks in automatic speech recognition (ASR), and is beneficial to integrate into multilingual end-to-end ASR as an additiona…
Streaming Align-Refine for Non-autoregressive Deliberation
Weiran Wang, Ke Hu, Tara N. Sainath
We propose a streaming non-autoregressive (non-AR) decoding algorithm to deliberate the hypothesis alignment of a streaming RNN-T model. Our algorithm facilitates a simple greedy d…
Improving the fusion of acoustic and text representations in RNN-T
Chao Zhang, Bo Li, Zhiyun Lu +2
The recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR). To estimate the output dist…