activity
20182026
most citedNeural Codec Language Models are Zero-Shot Text to Speech Synthesizers

163 citations · 299 across the 22 of their papers we have counts for

collaborators
Showing 2022Show all

7 papers · 1 filter

eess.AS2022★ 43 cited

BEATs: Audio Pre-Training with Acoustic Tokenizers

Sanyuan Chen, Yu Wu, Chengyi Wang +4

The massive growth of self-supervised learning (SSL) has been witnessed in language, vision, speech, and audio domains over the past few years. While discrete label prediction is w…

cs.SD2022★ 3 cited

TESSP: Text-Enhanced Self-Supervised Speech Pre-training

Zhuoyuan Yao, Shuo Ren, Sanyuan Chen +3

Self-supervised speech pre-training empowers the model with the contextual structure inherent in the speech signal while self-supervised text pre-training empowers the model with l…

eess.AS2022

Exploring WavLM on Speech Enhancement

Hyungchan Song, Sanyuan Chen, Zhuo Chen +5

There is a surge in interest in self-supervised learning approaches for end-to-end speech encoding in recent years as they have achieved great success. Especially, WavLM showed sta…

cs.CL2022★ 13 cited

SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data

Ziqiang Zhang, Sanyuan Chen, Long Zhou +8

How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this…

cs.CL2022

Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training

Chengyi Wang, Yiming Wang, Yu Wu +4

Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition. It usually requires a codebook obtained in an unsupe…

eess.AS2022★ 15 cited

Ultra Fast Speech Separation Model with Teacher Student Learning

Sanyuan Chen, Yu Wu, Zhuo Chen +5

Transformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer ten…