4 papers
ADEPT: A Dataset for Evaluating Prosody Transfer
Alexandra Torresquintero, Tian Huey Teh, Christopher G. R. Wallis +6
Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a referen…
Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis
Devang S Ram Mohan, Vivian Hu, Tian Huey Teh +6
Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to r…
Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning
Devang S Ram Mohan, Raphael Lenain, Lorenzo Foglianti +4
Modern approaches to text to speech require the entire input character sequence to be processed before any audio is synthesised. This latency limits the suitability of such models…
Phonological Features for 0-shot Multilingual Speech Synthesis
Marlene Staib, Tian Huey Teh, Alexandra Torresquintero +4
Code-switching---the intra-utterance use of multiple languages---is prevalent across the world. Within text-to-speech (TTS), multilingual models have been found to enable code-swit…