activity
20172022
most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 463 across the 19 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2022

Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification

Chao Zhang, Bo Li, Tara Sainath +4

Language identification is critical for many downstream tasks in automatic speech recognition (ASR), and is beneficial to integrate into multilingual end-to-end ASR as an additiona…

eess.AS2022

Improving the fusion of acoustic and text representations in RNN-T

Chao Zhang, Bo Li, Zhiyun Lu +2

The recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR). To estimate the output dist…

eess.AS2021

Residual Energy-Based Models for End-to-End Speech Recognition

Qiujia Li, Yu Zhang, Bo Li +2

End-to-end models with auto-regressive decoders have shown impressive results for automatic speech recognition (ASR). These models formulate the sequence-level probability as a pro…

eess.AS2020

Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition

Qiujia Li, David Qiu, Yu Zhang +5

For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based…

eess.AS2020

FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization

Jiahui Yu, Chung-Cheng Chiu, Bo Li +8

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible. However, emitting fast without degrading quality, as measure…

eess.AS2018

Bytes are All You Need: End-to-End Multilingual Speech Recognition and Synthesis with Bytes

Bo Li, Yu Zhang, Tara Sainath +2

We present two end-to-end models: Audio-to-Byte (A2B) and Byte-to-Audio (B2A), for multilingual speech recognition and synthesis. Prior work has predominantly used characters, sub-…