activity
20192021
most citedMulti-Speaker End-to-End Speech Synthesis

25 citations · 28 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG20213 cited

Topological Regularization for Graph Neural Networks Augmentation

Rui Song, Fausto Giunchiglia, Ke Zhao +1

The complexity and non-Euclidean structure of graph data hinder the development of data augmentation methods similar to those in computer vision. In this paper, we propose a featur…

eess.AS2020

DiffWave: A Versatile Diffusion Model for Audio Synthesis

Zhifeng Kong, Wei Ping, Jiaji Huang +2

In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts th…

math.FA2020

A property in vector-valued function spaces

Kexin Zhao, Dongni Tan

This paper deals with a property which is equivalent to generalised-lushness for separable spaces. It thus may be seemed as a geometrical property of a Banach space which ensures t…

cs.SD2019

WaveFlow: A Compact Flow-based Model for Raw Audio

Wei Ping, Kainan Peng, Kexin Zhao +1

In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood. It handles the long-range structure of 1-D wa…

cs.CL201925 cited

Multi-Speaker End-to-End Speech Synthesis

Jihyun Park, Kexin Zhao, Kainan Peng +1

In this work, we extend ClariNet (Ping et al., 2019), a fully end-to-end speech synthesis model (i.e., text-to-wave), to generate high-fidelity speech from multiple speakers. To mo…

cs.CL2019

Non-Autoregressive Neural Text-to-Speech

Kainan Peng, Wei Ping, Zhao Song +1

In this work, we propose ParaNet, a non-autoregressive seq2seq model that converts text to spectrogram. It is fully convolutional and brings 46.7 times speed-up over the lightweigh…