44 citations · 86 across the 5 of their papers we have counts for
13 papers
Speaker Generation
Daisy Stanton, Matt Shannon, Soroosh Mariooryad +4
This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs compe…
Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis
Ron J. Weiss, RJ Skerry-Ryan, Eric Battenberg +2
We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs. The architecture extends the Tacotron model by incorporating a normali…
Non-saturating GAN training as divergence minimization
Matt Shannon, Ben Poole, Soroosh Mariooryad +5
Non-saturating generative adversarial network (GAN) training is widely used and has continued to obtain groundbreaking results. However so far this approach has lacked strong theor…
Semi-Supervised Generative Modeling for Controllable Speech Synthesis
Raza Habib, Soroosh Mariooryad, Matt Shannon +5
We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…
Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad +4
Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in f…
Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
Yu Zhang, Ron J. Weiss, Heiga Zen +6
We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages. Moreover, the mode…