15 citations · 50 across the 18 of their papers we have counts for
19 papers
Exploring WavLM on Speech Enhancement
Hyungchan Song, Sanyuan Chen, Zhuo Chen +5
There is a surge in interest in self-supervised learning approaches for end-to-end speech encoding in recent years as they have achieved great success. Especially, WavLM showed sta…
LongFNT: Long-form Speech Recognition with Factorized Neural Transducer
Xun Gong, Yu Wu, Jinyu Li +4
Traditional automatic speech recognition~(ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is mor…
Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
Kun Wei, Long Zhou, Ziqiang Zhang +5
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity probl…
Robust Data2vec: Noise-robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
Qiu-Shi Zhu, Long Zhou, Jie Zhang +3
Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (A…
SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
Ziqiang Zhang, Long Zhou, Junyi Ao +4
The rapid development of single-modal pre-training has prompted researchers to pay more attention to cross-modal pre-training methods. In this paper, we propose a unified-modal spe…
Ultra Fast Speech Separation Model with Teacher Student Learning
Sanyuan Chen, Yu Wu, Zhuo Chen +5
Transformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer ten…