4 citations · 5 across the 6 of their papers we have counts for
8 papers
MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis
Yi Lei, Shan Yang, Xinsheng Wang +1
Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in…
Controllable Emotion Transfer For End-to-End Speech Synthesis
Tao Li, Shan Yang, Liumeng Xue +1
Emotion embedding space learned from references is a straightforward approach for emotion transfer in encoder-decoder structured emotional text to speech (TTS) systems. However, th…
Accent and Speaker Disentanglement in Many-to-many Voice Conversion
Zhichao Wang, Wenshuo Ge, Xiong Wang +6
This paper proposes an interesting voice and accent joint conversion approach, which can convert an arbitrary source speaker's voice to a target speaker with non-native accent. Thi…
Fine-grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis
Yi Lei, Shan Yang, Lei Xie
This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotion…
Learn2Sing: Target Speaker Singing Voice Synthesis by learning from a Singing Teacher
Heyang Xue, Shan Yang, Yi Lei +2
Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produc…
Exploiting Deep Sentential Context for Expressive End-to-End Speech Synthesis
Fengyu Yang, Shan Yang, Qinghua Wu +2
Attention-based seq2seq text-to-speech systems, especially those use self-attention networks (SAN), have achieved state-of-art performance. But an expressive corpus with rich proso…