activity
20182021
most citedSemi-Supervised Generative Modeling for Controllable Speech Synthesis

15 citations · 16 across the 3 of their papers we have counts for

collaborators

7 papers

cs.SD20211 cited

Speaker Generation

Daisy Stanton, Matt Shannon, Soroosh Mariooryad +4

This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs compe…

cs.LG2020

Non-saturating GAN training as divergence minimization

Matt Shannon, Ben Poole, Soroosh Mariooryad +5

Non-saturating generative adversarial network (GAN) training is widely used and has continued to obtain groundbreaking results. However so far this approach has lacked strong theor…

cs.CL201915 cited

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

Raza Habib, Soroosh Mariooryad, Matt Shannon +5

We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…

cs.CL2019

Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad +4

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in f…

cs.LG2019

Complex Evolution Recurrent Neural Networks (ceRNNs)

Izhak Shafran, Tom Bagby, R. J. Skerry-Ryan

Unitary Evolution Recurrent Neural Networks (uRNNs) have three attractive properties: (a) the unitary property, (b) the complex-valued nature, and (c) their efficient linear operat…

cs.CL2019

Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis

Eric Battenberg, Soroosh Mariooryad, Daisy Stanton +4

Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a…