activity
20182022
most citedMinimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition

16 citations · 101 across the 22 of their papers we have counts for

collaborators

27 papers

cs.SD20221 cited

NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS

Dongchao Yang, Songxiang Liu, Jianwei Yu +3

Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…

cs.SD20226 cited

The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022

Xiaoyi Qin, Na Li, Yuke Lin +4

This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For…

cs.SD2022

Improving Target Sound Extraction with Timestamp Information

Helin Wang, Dongchao Yang, Chao Weng +2

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…

eess.AS2022

The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

Naijun Zheng, Na Li, Xixin Wu +6

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in…

cs.CL202213 cited

Improving Mandarin End-to-End Speech Recognition with Word N-gram Language Model

Jinchuan Tian, Jianwei Yu, Chao Weng +2

Despite the rapid progress of end-to-end (E2E) automatic speech recognition (ASR), it has been shown that incorporating external language models (LMs) into the decoding can further…

cs.SD20213 cited

Simple Attention Module based Speaker Verification with Iterative noisy label detection

Xiaoyi Qin, Na Li, Chao Weng +2

Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speak…