activity
20192022
most citedSemi-Supervised Generative Modeling for Controllable Speech Synthesis

15 citations · 16 across the 3 of their papers we have counts for

collaborators

7 papers

cs.LG2022

Learning the joint distribution of two sequences using little or no paired data

Soroosh Mariooryad, Matt Shannon, Siyuan Ma +5

We present a noisy channel generative model of two sequences, for example text and speech, which enables uncovering the association between the two modalities when limited paired d…

cs.SD20211 cited

Speaker Generation

Daisy Stanton, Matt Shannon, Soroosh Mariooryad +4

This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs compe…

cs.CL2020

Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis

Ron J. Weiss, RJ Skerry-Ryan, Eric Battenberg +2

We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs. The architecture extends the Tacotron model by incorporating a normali…

cs.LG2020

Non-saturating GAN training as divergence minimization

Matt Shannon, Ben Poole, Soroosh Mariooryad +5

Non-saturating generative adversarial network (GAN) training is widely used and has continued to obtain groundbreaking results. However so far this approach has lacked strong theor…

cs.CL201915 cited

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

Raza Habib, Soroosh Mariooryad, Matt Shannon +5

We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…

cs.CL2019

Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad +4

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in f…