most citedMassive End-to-end Models for Short Search Queries

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS2024

Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models

Tsendsuren Munkhdalai, Youzheng Chen, Khe Chai Sim +3

Parameter efficient adaptation methods have become a key mechanism to train large pre-trained models for downstream tasks. However, their per-task parameter overhead is considered…

cs.CL2023

Audio-AdapterFusion: A Task-ID-free Approach for Efficient and Non-Destructive Multi-task Speech Recognition

Hillary Ngai, Rohan Agrawal, Neeraj Gaur +3

Adapters are an efficient, composable alternative to full fine-tuning of pre-trained models and help scale the deployment of large ASR models to many tasks. In practice, a task ID…

cs.CL2023

Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm

Weiran Wang, Zelin Wu, Diamantino Caseiro +10

Contextual biasing refers to the problem of biasing the automatic speech recognition (ASR) systems towards rare entities that are relevant to the specific user or application scena…

eess.AS20231 cited

Massive End-to-end Models for Short Search Queries

Weiran Wang, Rohit Prabhavalkar, Dongseong Hwang +11

In this work, we investigate two popular end-to-end automatic speech recognition (ASR) models, namely Connectionist Temporal Classification (CTC) and RNN-Transducer (RNN-T), for of…

eess.AS2023

Improving Speech Recognition for African American English With Audio Classification

Shefali Garg, Zhouyuan Huo, Khe Chai Sim +11

Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to m…

eess.AS2023

Modular Domain Adaptation for Conformer-Based Streaming ASR

Qiujia Li, Bo Li, Dongseong Hwang +2

Speech data from different domains has distinct acoustic and linguistic characteristics. It is common to train a single multidomain model such as a Conformer transducer for speech…