859 citations · 870 across the 5 of their papers we have counts for
7 papers
Learning Word-Level Confidence For Subword End-to-End ASR
David Qiu, Qiujia Li, Yanzhang He +9
We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed trainin…
A Better and Faster End-to-End Model for Streaming ASR
Bo Li, Anmol Gulati, Jiahui Yu +12
End-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by…
Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
Jiahui Yu, Wei Han, Anmol Gulati +5
Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible, while full-context ASR waits for the completion of a full sp…
Towards Fast and Accurate Streaming End-to-End ASR
Bo Li, Shuo-yiin Chang, Tara N. Sainath +4
End-to-end (E2E) models fold the acoustic, pronunciation and language models of a conventional speech recognition model into one neural network with a much smaller number of parame…
Improved Noisy Student Training for Automatic Speech Recognition
Daniel S. Park, Yu Zhang, Ye Jia +5
Recently, a semi-supervised learning method known as "noisy student training" has been shown to improve image classification performance of deep networks significantly. Noisy stude…
SpecAugment on Large Scale Datasets
Daniel S. Park, Yu Zhang, Chung-Cheng Chiu +5
Recently, SpecAugment, an augmentation scheme for automatic speech recognition that acts directly on the spectrogram of input utterances, has shown to be highly effective in enhanc…