3 citations · 6 across the 3 of their papers we have counts for
5 papers · 1 filter
Cloning one's voice using very limited data in the wild
Dongyang Dai, Yuanzhe Chen, Li Chen +6
With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily a…
Unsupervised Cross-Lingual Speech Emotion Recognition Using DomainAdversarial Neural Network
Xiong Cai, Zhiyong Wu, Kuo Zhong +3
By using deep learning approaches, Speech Emotion Recog-nition (SER) on a single domain has achieved many excellentresults. However, cross-domain SER is still a challenging taskdue…
Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition
Xiong Cai, Dongyang Dai, Zhiyong Wu +3
Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In…
Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams
Huirong Huang, Zhiyong Wu, Shiyin Kang +9
Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent meth…
Noise Robust TTS for Low Resource Speakers using Pre-trained Model and Speech Enhancement
Dongyang Dai, Li Chen, Yuping Wang +5
With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More a…