53 citations · 84 across the 7 of their papers we have counts for
Showing 2019Show all
3 papers · 1 filter
cs.SD2019
WaveFlow: A Compact Flow-based Model for Raw Audio
Wei Ping, Kainan Peng, Kexin Zhao +1
In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood. It handles the long-range structure of 1-D wa…
cs.CL2019★ 25 cited
Multi-Speaker End-to-End Speech Synthesis
Jihyun Park, Kexin Zhao, Kainan Peng +1
In this work, we extend ClariNet (Ping et al., 2019), a fully end-to-end speech synthesis model (i.e., text-to-wave), to generate high-fidelity speech from multiple speakers. To mo…
cs.CL2019
Non-Autoregressive Neural Text-to-Speech
Kainan Peng, Wei Ping, Zhao Song +1
In this work, we propose ParaNet, a non-autoregressive seq2seq model that converts text to spectrogram. It is fully convolutional and brings 46.7 times speed-up over the lightweigh…