387 citations · 764 across the 22 of their papers we have counts for
16 papers · 1 filter
A Better and Faster End-to-End Model for Streaming ASR
Bo Li, Anmol Gulati, Jiahui Yu +12
End-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by…
Cascaded encoders for unifying streaming and non-streaming ASR
Arun Narayanan, Tara N. Sainath, Ruoming Pang +5
End-to-end (E2E) automatic speech recognition (ASR) models, by now, have shown competitive performance on several benchmarks. These models are structured to either operate in strea…
Unsupervised Learning of Disentangled Speech Content and Style Representation
Andros Tjandra, Ruoming Pang, Yu Zhang +1
We present an approach for unsupervised learning of speech representation disentangling contents and styles. Our model consists of: (1) a local encoder that captures per-frame info…
Improving Streaming Automatic Speech Recognition With Non-Streaming Model Distillation On Unsupervised Data
Thibault Doutre, Wei Han, Min Ma +7
Streaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech wi…
FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
Jiahui Yu, Chung-Cheng Chiu, Bo Li +8
Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible. However, emitting fast without degrading quality, as measure…
Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
Yu Zhang, James Qin, Daniel S. Park +5
We employ a combination of recent developments in semi-supervised learning for automatic speech recognition to obtain state-of-the-art results on LibriSpeech utilizing the unlabele…