27 citations · 92 across the 15 of their papers we have counts for
8 papers · 1 filter
Input Length Matters: Improving RNN-T and MWER Training for Long-form Telephony Speech Recognition
Zhiyun Lu, Yanwei Pan, Thibault Doutre +5
End-to-end models have achieved state-of-the-art results on several automatic speech recognition tasks. However, they perform poorly when evaluated on long-form data, e.g., minutes…
Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition
Qiujia Li, Yu Zhang, David Qiu +3
As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research…
BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
Yu Zhang, Daniel S. Park, Wei Han +23
We summarize the results of a host of efforts using giant automatic speech recognition (ASR) models pre-trained using large, diverse unlabeled datasets containing approximately a m…
Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
David Qiu, Yanzhang He, Qiujia Li +3
Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utter…
Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models
Thibault Doutre, Wei Han, Chung-Cheng Chiu +3
Streaming end-to-end automatic speech recognition (ASR) systems are widely used in everyday applications that require transcribing speech to text in real-time. Their minimal latenc…
Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models
Zhiyun Lu, Wei Han, Yu Zhang +1
Although end-to-end automatic speech recognition (e2e ASR) models are widely deployed in many applications, there have been very few studies to understand models' robustness agains…