Publications (13)
Semi-Supervised Generative Modeling for Controllable Speech Synthesis
Raza Habib, Soroosh Mariooryad, Matt Shannon +5
We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…
Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis
Daisy Stanton, Yuxuan Wang, RJ Skerry-Ryan
Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-ar…
Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad +4
Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in f…
Uncovering Latent Style Factors for Expressive Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Ying Xiao +5
Prosodic modeling is a core problem in speech synthesis. The key challenge is producing desirable prosody from textual input containing only phonetic information. In this prelimina…
Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton +11
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these component…
Speaker Generation
Daisy Stanton, Matt Shannon, Soroosh Mariooryad +4
This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs compe…