activity
20192021
most citedDeep Learning for Audio Signal Processing

859 citations · 870 across the 5 of their papers we have counts for

collaborators

7 papers

eess.AS20212 cited

Learning Word-Level Confidence For Subword End-to-End ASR

David Qiu, Qiujia Li, Yanzhang He +9

We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed trainin…

eess.AS20202 cited

A Better and Faster End-to-End Model for Streaming ASR

Bo Li, Anmol Gulati, Jiahui Yu +12

End-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by…

cs.CL2020

Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling

Jiahui Yu, Wei Han, Anmol Gulati +5

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible, while full-context ASR waits for the completion of a full sp…

eess.AS20203 cited

Towards Fast and Accurate Streaming End-to-End ASR

Bo Li, Shuo-yiin Chang, Tara N. Sainath +4

End-to-end (E2E) models fold the acoustic, pronunciation and language models of a conventional speech recognition model into one neural network with a much smaller number of parame…

eess.AS2020

Improved Noisy Student Training for Automatic Speech Recognition

Daniel S. Park, Yu Zhang, Ye Jia +5

Recently, a semi-supervised learning method known as "noisy student training" has been shown to improve image classification performance of deep networks significantly. Noisy stude…

eess.AS20194 cited

SpecAugment on Large Scale Datasets

Daniel S. Park, Yu Zhang, Chung-Cheng Chiu +5

Recently, SpecAugment, an augmentation scheme for automatic speech recognition that acts directly on the spectrogram of input utterances, has shown to be highly effective in enhanc…