activity
20202022
most citedDomain segmentation and adjustment for generalized zero-shot learning

6 citations · 12 across the 8 of their papers we have counts for

collaborators

8 papers

cs.SD2022

UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis

Yi Lei, Shan Yang, Xinsheng Wang +4

Text-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying…

cs.SD2022

Learn2Sing 2.0: Diffusion and Mutual Information-Based Target Speaker SVS by Learning from Singing Teacher

Heyang Xue, Xinsheng Wang, Yongmao Zhang +3

Building a high-quality singing corpus for a person who is not good at singing is non-trivial, thus making it challenging to create a singing voice synthesizer for this person. Lea…

cs.SD20221 cited

Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis

Yu Wang, Xinsheng Wang, Pengcheng Zhu +6

This paper introduces Opencpop, a publicly available high-quality Mandarin singing corpus designed for singing voice synthesis (SVS). The corpus consists of 100 popular Mandarin so…

cs.SD20224 cited

MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis

Yi Lei, Shan Yang, Xinsheng Wang +1

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in…

cs.CV2021

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

Xinsheng Wang, Qicong Xie, Jihua Zhu +2

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. I…

cs.CV2020

Show and Speak: Directly Synthesize Spoken Description of Images

Xinsheng Wang, Siyuan Feng, Jihua Zhu +2

This paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the…