2 citations · 3 across the 6 of their papers we have counts for
6 papers
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
Rohit Prabhavalkar, Zhong Meng, Weiran Wang +7
The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. W…
Massive End-to-end Models for Short Search Queries
Weiran Wang, Rohit Prabhavalkar, Dongseong Hwang +11
In this work, we investigate two popular end-to-end automatic speech recognition (ASR) models, namely Connectionist Temporal Classification (CTC) and RNN-Transducer (RNN-T), for of…
Improving Speech Recognition for African American English With Audio Classification
Shefali Garg, Zhouyuan Huo, Khe Chai Sim +11
Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to m…
Edit Distance based RL for RNNT decoding
Dongseong Hwang, Changwan Ryu, Khe Chai Sim
RNN-T is currently considered the industry standard in ASR due to its exceptional WERs in various benchmark tests and its ability to support seamless streaming and longform transcr…
Modular Domain Adaptation for Conformer-Based Streaming ASR
Qiujia Li, Bo Li, Dongseong Hwang +2
Speech data from different domains has distinct acoustic and linguistic characteristics. It is common to train a single multidomain model such as a Conformer transducer for speech…
Efficient Domain Adaptation for Speech Foundation Models
Bo Li, Dongseong Hwang, Zhouyuan Huo +8
Foundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benef…