13 citations · 46 across the 11 of their papers we have counts for
5 papers · 1 filter
Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
Kun Wei, Long Zhou, Ziqiang Zhang +5
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity probl…
Cross-Lingual Text-to-Speech Using Multi-Task Learning and Speaker Classifier Joint Training
J. Yang, Lei He
In cross-lingual speech synthesis, the speech in various languages can be synthesized for a monoglot speaker. Normally, only the data of monoglot speakers are available for model t…
Cross-speaker Style Transfer with Prosody Bottleneck in Neural Speech Synthesis
Shifeng Pan, Lei He
Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expres…
Conversational End-to-End TTS for Voice Agent
Haohan Guo, Shaofei Zhang, Frank K. Soong +2
End-to-end neural TTS has achieved superior performance on reading style speech synthesis. However, it's still a challenge to build a high-quality conversational TTS due to the lim…
Feature reinforcement with word embedding and parsing information in neural TTS
Huaiping Ming, Lei He, Haohan Guo +1
In this paper, we propose a feature reinforcement method under the sequence-to-sequence neural text-to-speech (TTS) synthesis framework. The proposed method utilizes the multiple i…