most citedMsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis

4 citations · 5 across the 7 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2022

UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis

Yi Lei, Shan Yang, Xinsheng Wang +4

Text-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying…

cs.SD20224 cited

MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis

Yi Lei, Shan Yang, Xinsheng Wang +1

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in…

cs.SD2020

Controllable Emotion Transfer For End-to-End Speech Synthesis

Tao Li, Shan Yang, Liumeng Xue +1

Emotion embedding space learned from references is a straightforward approach for emotion transfer in encoder-decoder structured emotional text to speech (TTS) systems. However, th…

cs.SD2020

Accent and Speaker Disentanglement in Many-to-many Voice Conversion

Zhichao Wang, Wenshuo Ge, Xiong Wang +6

This paper proposes an interesting voice and accent joint conversion approach, which can convert an arbitrary source speaker's voice to a target speaker with non-native accent. Thi…

cs.SD2020

Fine-grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis

Yi Lei, Shan Yang, Lei Xie

This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotion…

cs.SD20201 cited

Learn2Sing: Target Speaker Singing Voice Synthesis by learning from a Singing Teacher

Heyang Xue, Shan Yang, Yi Lei +2

Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produc…