activity
20182022
most citedFlowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis

81 citations · 105 across the 5 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2021

One TTS Alignment To Rule Them All

Rohan Badlani, Adrian Łancucki, Kevin J. Shih +3

Speech-to-text alignment is a critical component of neural textto-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-l…

cs.SD202081 cited

Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis

Rafael Valle, Kevin Shih, Ryan Prenger +1

In this paper we propose Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis with control over speech variation and style transfer. Flowtron borr…

cs.SD2019

Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens

Rafael Valle, Jason Li, Ryan Prenger +1

Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data. By explicitly conditioning…

cs.SD2018

WaveGlow: A Flow-based Generative Network for Speech Synthesis

Ryan Prenger, Rafael Valle, Bryan Catanzaro

In this paper we propose WaveGlow: a flow-based network capable of generating high quality speech from mel-spectrograms. WaveGlow combines insights from Glow and WaveNet in order t…

cs.SD201820 cited

Attacking Speaker Recognition With Deep Generative Models

Wilson Cai, Anish Doshi, Rafael Valle

In this paper we investigate the ability of generative adversarial networks (GANs) to synthesize spoofing attacks on modern speaker recognition systems. We first show that samples…