58 citations · 97 across the 38 of their papers we have counts for
15 papers · 1 filter
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
Yimin Deng, Jianzong Wang, Xulong Zhang +2
Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are…
Semi-Supervised Learning Based on Reference Model for Low-resource TTS
Xulong Zhang, Jianzong Wang, Ning Cheng +1
Most previous neural text-to-speech (TTS) methods are mainly based on supervised learning methods, which means they depend on a large training dataset and hard to achieve comparabl…
MetaSpeech: Speech Effects Switch Along with Environment for Metaverse
Xulong Zhang, Jianzong Wang, Ning Cheng +1
Metaverse expands the physical world to a new dimension, and the physical environment and Metaverse environment can be directly connected and entered. Voice is an indispensable com…
Improving Speech Representation Learning via Speech-level and Phoneme-level Masking Approach
Xulong Zhang, Jianzong Wang, Ning Cheng +2
Recovering the masked speech frames is widely applied in speech representation learning. However, most of these models use random masking in the pre-training. In this work, we prop…
Adapitch: Adaption Multi-Speaker Text-to-Speech Conditioned on Pitch Disentangling with Untranscribed Data
Xulong Zhang, Jianzong Wang, Ning Cheng +1
In this paper, we proposed Adapitch, a multi-speaker TTS method that makes adaptation of the supervised module with untranscribed data. We design two self supervised modules to tra…
Speech Augmentation Based Unsupervised Learning for Keyword Spotting
Jian Luo, Jianzong Wang, Ning Cheng +2
In this paper, we investigated a speech augmentation based unsupervised learning approach for keyword spotting (KWS) task. KWS is a useful speech application, yet also heavily depe…