activity
20202022
most citedUnified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition

46 citations · 84 across the 9 of their papers we have counts for

collaborators

9 papers

cs.SD2022★ 3 cited

TESSP: Text-Enhanced Self-Supervised Speech Pre-training

Zhuoyuan Yao, Shuo Ren, Sanyuan Chen +3

Self-supervised speech pre-training empowers the model with the contextual structure inherent in the speech signal while self-supervised text pre-training empowers the model with l…

cs.CL2022★ 13 cited

SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data

Ziqiang Zhang, Sanyuan Chen, Long Zhou +8

How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this…

cs.SD2022★ 2 cited

WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit

Binbin Zhang, Di Wu, Zhendong Peng +7

Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address…

cs.SD2021★ 2 cited

Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR

Fan Yu, Haoneng Luo, Pengcheng Guo +6

Continuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with com…

cs.SD2021★ 16 cited

WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit

Zhuoyuan Yao, Di Wu, Xiong Wang +7

In this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify st…

cs.SD2020★ 46 cited

Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition

Binbin Zhang, Di Wu, Zhuoyuan Yao +7

In this paper, we present a novel two-pass approach to unify streaming and non-streaming end-to-end (E2E) speech recognition in a single model. Our model adopts the hybrid CTC/atte…