25 citations · 25 across the 1 of their papers we have counts for
6 papers
WaveFlow: A Compact Flow-based Model for Raw Audio
Wei Ping, Kainan Peng, Kexin Zhao +1
In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood. It handles the long-range structure of 1-D wa…
Incremental Text-to-Speech Synthesis with Prefix-to-Prefix Framework
Mingbo Ma, Baigong Zheng, Kaibo Liu +5
Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness. However, these efforts s…
Multi-Speaker End-to-End Speech Synthesis
Jihyun Park, Kexin Zhao, Kainan Peng +1
In this work, we extend ClariNet (Ping et al., 2019), a fully end-to-end speech synthesis model (i.e., text-to-wave), to generate high-fidelity speech from multiple speakers. To mo…
Non-Autoregressive Neural Text-to-Speech
Kainan Peng, Wei Ping, Zhao Song +1
In this work, we propose ParaNet, a non-autoregressive seq2seq model that converts text to spectrogram. It is fully convolutional and brings 46.7 times speed-up over the lightweigh…
ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
Wei Ping, Kainan Peng, Jitong Chen
In this work, we propose a new solution for parallel wave generation by WaveNet. In contrast to parallel WaveNet (van den Oord et al., 2018), we distill a Gaussian inverse autoregr…
Neural Voice Cloning with a Few Samples
Sercan O. Arik, Jitong Chen, Kainan Peng +2
Voice cloning is a highly desired feature for personalized speech interfaces. Neural network based speech synthesis has been shown to generate high quality speech for a large numbe…