20 citations · 129 across the 28 of their papers we have counts for
12 papers · 1 filter
HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Dongchao Yang, Songxiang Liu, Rongjie Huang +3
Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations. Nowadays, audio codec models are increasingly…
Diverse and Expressive Speech Prosody Prediction with Denoising Diffusion Probabilistic Model
Xiang Li, Songxiang Liu, Max W. Y. Lam +3
Expressive human speech generally abounds with rich and flexible speech prosody variations. The speech prosody predictors in existing expressive speech synthesis methods mostly pro…
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu +3
Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…
The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022
Xiaoyi Qin, Na Li, Yuke Lin +4
This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For…
Improving Target Sound Extraction with Timestamp Information
Helin Wang, Dongchao Yang, Chao Weng +2
Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…
Simple Attention Module based Speaker Verification with Iterative noisy label detection
Xiaoyi Qin, Na Li, Chao Weng +2
Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speak…