activity
20162022
most citedDeep Learning for Audio Signal Processing

859 citations · 1.1k across the 26 of their papers we have counts for

collaborators

43 papers

cs.LG2022

Resource-Efficient Transfer Learning From Speech Foundation Model Using Hierarchical Feature Fusion

Zhouyuan Huo, Khe Chai Sim, Bo Li +3

Self-supervised pre-training of a speech foundation model, followed by supervised fine-tuning, has shown impressive quality improvements on automatic speech recognition (ASR) tasks…

cs.CL20221 cited

JOIST: A Joint Speech and Text Streaming Model For ASR

Tara N. Sainath, Rohit Prabhavalkar, Ankur Bapna +6

We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous wo…

cs.CL2022

Scaling Up Deliberation for Multilingual ASR

Ke Hu, Bo Li, Tara N. Sainath

Multilingual end-to-end automatic speech recognition models are attractive due to its simplicity in training and deployment. Recent work on large-scale training of such models has…

eess.AS2022

Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification

Chao Zhang, Bo Li, Tara Sainath +4

Language identification is critical for many downstream tasks in automatic speech recognition (ASR), and is beneficial to integrate into multilingual end-to-end ASR as an additiona…

cs.CL2022

Streaming Align-Refine for Non-autoregressive Deliberation

Weiran Wang, Ke Hu, Tara N. Sainath

We propose a streaming non-autoregressive (non-AR) decoding algorithm to deliberate the hypothesis alignment of a streaming RNN-T model. Our algorithm facilitates a simple greedy d…

eess.AS2022

Improving the fusion of acoustic and text representations in RNN-T

Chao Zhang, Bo Li, Zhiyun Lu +2

The recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR). To estimate the output dist…