activity
20182022
most citedDeep Learning for Audio Signal Processing

859 citations · 867 across the 7 of their papers we have counts for

collaborators

8 papers

eess.AS20202 cited

A Better and Faster End-to-End Model for Streaming ASR

Bo Li, Anmol Gulati, Jiahui Yu +12

End-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by…

eess.AS2020

FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization

Jiahui Yu, Chung-Cheng Chiu, Bo Li +8

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible. However, emitting fast without degrading quality, as measure…

eess.AS20203 cited

Towards Fast and Accurate Streaming End-to-End ASR

Bo Li, Shuo-yiin Chang, Tara N. Sainath +4

End-to-end (E2E) models fold the acoustic, pronunciation and language models of a conventional speech recognition model into one neural network with a much smaller number of parame…

cs.CL20203 cited

A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency

Tara N. Sainath, Yanzhang He, Bo Li +26

Thus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e…

eess.AS2019

On Neural Phone Recognition of Mixed-Source ECoG Signals

Ahmed Hussen Abdelaziz, Shuo-Yiin Chang, Nelson Morgan +5

The emerging field of neural speech recognition (NSR) using electrocorticography has recently attracted remarkable research interest for studying how human brains recognize speech…

eess.AS2019

Personal VAD: Speaker-Conditioned Voice Activity Detection

Shaojin Ding, Quan Wang, Shuo-yiin Chang +2

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming o…