46 citations · 84 across the 9 of their papers we have counts for
9 papers
TESSP: Text-Enhanced Self-Supervised Speech Pre-training
Zhuoyuan Yao, Shuo Ren, Sanyuan Chen +3
Self-supervised speech pre-training empowers the model with the contextual structure inherent in the speech signal while self-supervised text pre-training empowers the model with l…
SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
Ziqiang Zhang, Sanyuan Chen, Long Zhou +8
How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this…
WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit
Binbin Zhang, Di Wu, Zhendong Peng +7
Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address…
Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
Fan Yu, Haoneng Luo, Pengcheng Guo +6
Continuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with com…
WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
Zhuoyuan Yao, Di Wu, Xiong Wang +7
In this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify st…
Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
Binbin Zhang, Di Wu, Zhuoyuan Yao +7
In this paper, we present a novel two-pass approach to unify streaming and non-streaming end-to-end (E2E) speech recognition in a single model. Our model adopts the hybrid CTC/atte…