Multi-reference Tacotron by Intercross Training for Style Disentangling,Transfer and Control in Speech Synthesis
arXiv:1904.02373
Abstract
Speech style control and transfer techniques aim to enrich the diversity and expressiveness of synthesized speech. Existing approaches model all speech styles into one representation, lacking the ability to control a specific speech feature independently. To address this issue, we introduce a novel multi-reference structure to Tacotron and propose intercross training approach, which together ensure that each sub-encoder of the multi-reference encoder independently disentangles and controls a specific style. Experimental results show that our model is able to control and transfer desired speech styles individually.
Submitted for Interspeech 2019, 5 pages
References in corpus (1)
Cited by in corpus (8)
- Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech
- Multi-Reference Neural TTS Stylization with Adversarial Cycle Consistency
- GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis
- Improving Performance of Seen and Unseen Speech Style Transfer in End-to-end Neural TTS
- Controllable Emotion Transfer For End-to-End Speech Synthesis
- Fine-grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis
- Emotional speech synthesis with rich and granularized control
- Controllable Context-aware Conversational Speech Synthesis