3 papers
cs.SD2022
Multi-Speaker Multi-Style Speech Synthesis with Timbre and Style Disentanglement
Wei Song, Yanghao Yue, Ya-jie Zhang +3
Disentanglement of a speaker's timbre and style is very important for style transfer in multi-speaker multi-style text-to-speech (TTS) scenarios. With the disentanglement of timbre…
eess.AS2020
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis
Guanghui Xu, Wei Song, Zhengchen Zhang +3
Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which ma…
cs.CL2019
Building a mixed-lingual neural TTS system with only monolingual data
Liumeng Xue, Wei Song, Guanghui Xu +2
When deploying a Chinese neural text-to-speech (TTS) synthesis system, one of the challenges is to synthesize Chinese utterances with English phrases or words embedded. This paper…