16 citations · 101 across the 22 of their papers we have counts for
27 papers
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu +3
Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…
The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022
Xiaoyi Qin, Na Li, Yuke Lin +4
This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For…
Improving Target Sound Extraction with Timestamp Information
Helin Wang, Dongchao Yang, Chao Weng +2
Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…
The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge
Naijun Zheng, Na Li, Xixin Wu +6
This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in…
Improving Mandarin End-to-End Speech Recognition with Word N-gram Language Model
Jinchuan Tian, Jianwei Yu, Chao Weng +2
Despite the rapid progress of end-to-end (E2E) automatic speech recognition (ASR), it has been shown that incorporating external language models (LMs) into the decoding can further…
Simple Attention Module based Speaker Verification with Iterative noisy label detection
Xiaoyi Qin, Na Li, Chao Weng +2
Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speak…