activity
20162022
most citedParallel WaveNet: Fast High-Fidelity Speech Synthesis

343 citations · 615 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20221 cited

WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration

Yuma Koizumi, Kohei Yatabe, Heiga Zen +1

Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characteriz…

eess.AS20215 cited

WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis

Nanxin Chen, Yu Zhang, Heiga Zen +4

This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density…

eess.AS2020

WaveGrad: Estimating Gradients for Waveform Generation

Nanxin Chen, Yu Zhang, Heiga Zen +3

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and di…

eess.AS202015 cited

Generating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody prior

Guangzhi Sun, Yu Zhang, Ron J. Weiss +5

Recent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-gr…

eess.AS20207 cited

Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis

Guangzhi Sun, Yu Zhang, Ron J. Weiss +3

This paper proposes a hierarchical, fine-grained and interpretable latent variable model for prosody based on the Tacotron 2 text-to-speech model. It achieves multi-resolution mode…