15 citations · 27 across the 5 of their papers we have counts for
9 papers
Learning the joint distribution of two sequences using little or no paired data
Soroosh Mariooryad, Matt Shannon, Siyuan Ma +5
We present a noisy channel generative model of two sequences, for example text and speech, which enables uncovering the association between the two modalities when limited paired d…
Global Normalization for Streaming Speech Recognition in a Modular Framework
Ehsan Variani, Ke Wu, Michael Riley +3
We introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact c…
Speaker Generation
Daisy Stanton, Matt Shannon, Soroosh Mariooryad +4
This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs compe…
Non-saturating GAN training as divergence minimization
Matt Shannon, Ben Poole, Soroosh Mariooryad +5
Non-saturating generative adversarial network (GAN) training is widely used and has continued to obtain groundbreaking results. However so far this approach has lacked strong theor…
Properties of f-divergences and f-GAN training
Matt Shannon
In this technical report we describe some properties of f-divergences and f-GAN training. We present an elementary derivation of the f-divergence lower bounds which form the basis…
Semi-Supervised Generative Modeling for Controllable Speech Synthesis
Raza Habib, Soroosh Mariooryad, Matt Shannon +5
We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…