4 papers
Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech
Sri Karlapati, Ammar Abbas, Zack Hodari +4
In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn…
CAMP: a Two-Stage Approach to Modelling Prosody in Context
Zack Hodari, Alexis Moinet, Sri Karlapati +6
Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody…
Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of F0
Zack Hodari, Catherine Lai, Simon King
In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However,…
Using generative modelling to produce varied intonation for speech synthesis
Zack Hodari, Oliver Watts, Simon King
Unlike human speakers, typical text-to-speech (TTS) systems are unable to produce multiple distinct renditions of a given sentence. This has previously been addressed by adding exp…