most citedEfficient Domain Adaptation for Speech Foundation Models

2 citations · 3 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models

Rohit Prabhavalkar, Zhong Meng, Weiran Wang +7

The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. W…

eess.AS20231 cited

Massive End-to-end Models for Short Search Queries

Weiran Wang, Rohit Prabhavalkar, Dongseong Hwang +11

In this work, we investigate two popular end-to-end automatic speech recognition (ASR) models, namely Connectionist Temporal Classification (CTC) and RNN-Transducer (RNN-T), for of…

eess.AS2023

Improving Speech Recognition for African American English With Audio Classification

Shefali Garg, Zhouyuan Huo, Khe Chai Sim +11

Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to m…

cs.SD2023

Edit Distance based RL for RNNT decoding

Dongseong Hwang, Changwan Ryu, Khe Chai Sim

RNN-T is currently considered the industry standard in ASR due to its exceptional WERs in various benchmark tests and its ability to support seamless streaming and longform transcr…

eess.AS2023

Modular Domain Adaptation for Conformer-Based Streaming ASR

Qiujia Li, Bo Li, Dongseong Hwang +2

Speech data from different domains has distinct acoustic and linguistic characteristics. It is common to train a single multidomain model such as a Conformer transducer for speech…

cs.CL20232 cited

Efficient Domain Adaptation for Speech Foundation Models

Bo Li, Dongseong Hwang, Zhouyuan Huo +8

Foundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benef…