4 citations · 4 across the 3 of their papers we have counts for
3 papers
Parallel WaveNet conditioned on VAE latent vectors
Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4
Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…
BOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian Optimization
Henry B. Moss, Vatsal Aggarwal, Nishant Prateek +2
We present BOFFIN TTS (Bayesian Optimization For FIne-tuning Neural Text To Speech), a novel approach for few-shot speaker adaptation. Here, the task is to fine-tune a pre-trained…
Using VAEs and Normalizing Flows for One-shot Text-To-Speech Synthesis of Expressive Speech
Vatsal Aggarwal, Marius Cotescu, Nishant Prateek +2
We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement…