3 papers
eess.AS2021
A learned conditional prior for the VAE acoustic space of a TTS system
Penny Karanasou, Sri Karlapati, Alexis Moinet +5
Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow mult…
eess.AS2020
Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech
Sri Karlapati, Ammar Abbas, Zack Hodari +4
In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn…
cs.CL2018
Selecting Machine-Translated Data for Quick Bootstrapping of a Natural Language Understanding System
Judith Gaspers, Penny Karanasou, Rajen Chatterjee
This paper investigates the use of Machine Translation (MT) to bootstrap a Natural Language Understanding (NLU) system for a new language for the use case of a large-scale voice-co…