7 citations · 7 across the 3 of their papers we have counts for
6 papers · 1 filter
Distribution augmentation for low-resource expressive text-to-speech
Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar +8
This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data.…
Multi-Scale Spectrogram Modelling for Neural Text-to-Speech
Ammar Abbas, Bajibabu Bollepalli, Alexis Moinet +6
We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody. We present a generic multi-scale spectrog…
A learned conditional prior for the VAE acoustic space of a TTS system
Penny Karanasou, Sri Karlapati, Alexis Moinet +5
Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow mult…
Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech
Sri Karlapati, Ammar Abbas, Zack Hodari +4
In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn…
CAMP: a Two-Stage Approach to Modelling Prosody in Context
Zack Hodari, Alexis Moinet, Sri Karlapati +6
Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody…
CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech
Sri Karlapati, Alexis Moinet, Arnaud Joly +3
Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects l…