859 citations · 867 across the 7 of their papers we have counts for
8 papers
A Better and Faster End-to-End Model for Streaming ASR
Bo Li, Anmol Gulati, Jiahui Yu +12
End-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by…
FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
Jiahui Yu, Chung-Cheng Chiu, Bo Li +8
Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible. However, emitting fast without degrading quality, as measure…
Towards Fast and Accurate Streaming End-to-End ASR
Bo Li, Shuo-yiin Chang, Tara N. Sainath +4
End-to-end (E2E) models fold the acoustic, pronunciation and language models of a conventional speech recognition model into one neural network with a much smaller number of parame…
A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency
Tara N. Sainath, Yanzhang He, Bo Li +26
Thus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e…
On Neural Phone Recognition of Mixed-Source ECoG Signals
Ahmed Hussen Abdelaziz, Shuo-Yiin Chang, Nelson Morgan +5
The emerging field of neural speech recognition (NSR) using electrocorticography has recently attracted remarkable research interest for studying how human brains recognize speech…
Personal VAD: Speaker-Conditioned Voice Activity Detection
Shaojin Ding, Quan Wang, Shuo-yiin Chang +2
In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming o…