papers

Publications (13)

cs.CL2019

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

Raza Habib, Soroosh Mariooryad, Matt Shannon +5

We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…

cs.CL2018

Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis

Daisy Stanton, Yuxuan Wang, RJ Skerry-Ryan

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-ar…

cs.CL2020

Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad +4

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in f…

cs.CL2017

Uncovering Latent Style Factors for Expressive Speech Synthesis

Yuxuan Wang, RJ Skerry-Ryan, Ying Xiao +5

Prosodic modeling is a core problem in speech synthesis. The key challenge is producing desirable prosody from textual input containing only phonetic information. In this prelimina…

cs.CL2017

Tacotron: Towards End-to-End Speech Synthesis

Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton +11

A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these component…

cs.SD2021

Speaker Generation

Daisy Stanton, Matt Shannon, Soroosh Mariooryad +4

This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs compe…