activity
20172022
most citedUncovering Latent Style Factors for Expressive Speech Synthesis

44 citations · 60 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL201915 cited

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

Raza Habib, Soroosh Mariooryad, Matt Shannon +5

We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial super…

cs.CL2019

Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad +4

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in f…

cs.CL2019

Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis

Eric Battenberg, Soroosh Mariooryad, Daisy Stanton +4

Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a…

cs.CL2018

Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis

Daisy Stanton, Yuxuan Wang, RJ Skerry-Ryan

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-ar…

cs.CL2018

Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

RJ Skerry-Ryan, Eric Battenberg, Ying Xiao +6

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing t…

cs.CL2018

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

Yuxuan Wang, Daisy Stanton, Yu Zhang +7

In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The emb…