163 citations · 299 across the 22 of their papers we have counts for
7 papers · 1 filter
BEATs: Audio Pre-Training with Acoustic Tokenizers
Sanyuan Chen, Yu Wu, Chengyi Wang +4
The massive growth of self-supervised learning (SSL) has been witnessed in language, vision, speech, and audio domains over the past few years. While discrete label prediction is w…
TESSP: Text-Enhanced Self-Supervised Speech Pre-training
Zhuoyuan Yao, Shuo Ren, Sanyuan Chen +3
Self-supervised speech pre-training empowers the model with the contextual structure inherent in the speech signal while self-supervised text pre-training empowers the model with l…
Exploring WavLM on Speech Enhancement
Hyungchan Song, Sanyuan Chen, Zhuo Chen +5
There is a surge in interest in self-supervised learning approaches for end-to-end speech encoding in recent years as they have achieved great success. Especially, WavLM showed sta…
SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
Ziqiang Zhang, Sanyuan Chen, Long Zhou +8
How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this…
Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training
Chengyi Wang, Yiming Wang, Yu Wu +4
Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition. It usually requires a codebook obtained in an unsupe…
Ultra Fast Speech Separation Model with Teacher Student Learning
Sanyuan Chen, Yu Wu, Zhuo Chen +5
Transformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer ten…