3 citations · 17 across the 15 of their papers we have counts for
4 papers · 1 filter
MnTTS2: An Open-Source Multi-Speaker Mongolian Text-to-Speech Synthesis Dataset
Kailin Liang, Bin Liu, Yifan Hu +3
Text-to-Speech (TTS) synthesis for low-resource languages is an attractive research issue in academia and industry nowadays. Mongolian is the official language of the Inner Mongoli…
Decoupling Speaker-Independent Emotions for Voice Conversion Via Source-Filter Networks
Zhaojie Luo, Shoufeng Lin, Rui Liu +3
Emotional voice conversion (VC) aims to convert a neutral voice to an emotional (e.g. happy) one while retaining the linguistic information and speaker identity. We note that the d…
Modeling Prosodic Phrasing with Multi-Task Learning in Tacotron-based TTS
Rui Liu, Berrak Sisman, Feilong Bao +2
Tacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for lo…
WaveTTS: Tacotron-based TTS with Joint Time-Frequency Domain Loss
Rui Liu, Berrak Sisman, Feilong Bao +2
Tacotron-based text-to-speech (TTS) systems directly synthesize speech from text input. Such frameworks typically consist of a feature prediction network that maps character sequen…