most citedSpeaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

3 citations · 6 across the 3 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2021

Cloning one's voice using very limited data in the wild

Dongyang Dai, Yuanzhe Chen, Li Chen +6

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily a…

eess.AS20203 cited

Unsupervised Cross-Lingual Speech Emotion Recognition Using DomainAdversarial Neural Network

Xiong Cai, Zhiyong Wu, Kuo Zhong +3

By using deep learning approaches, Speech Emotion Recog-nition (SER) on a single domain has achieved many excellentresults. However, cross-domain SER is still a challenging taskdue…

eess.AS2020

Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

Xiong Cai, Dongyang Dai, Zhiyong Wu +3

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In…

eess.AS20203 cited

Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

Huirong Huang, Zhiyong Wu, Shiyin Kang +9

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent meth…

eess.AS2020

Noise Robust TTS for Low Resource Speakers using Pre-trained Model and Speech Enhancement

Dongyang Dai, Li Chen, Yuping Wang +5

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More a…