35 citations · 42 across the 4 of their papers we have counts for
4 papers
Robust MelGAN: A robust universal neural vocoder for high-fidelity TTS
Kun Song, Jian Cong, Xinsheng Wang +4
In current two-stage neural text-to-speech (TTS) paradigm, it is ideal to have a universal neural vocoder, once trained, which is robust to imperfect mel-spectrogram predicted from…
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
Xu Tan, Jiawei Chen, Haohe Liu +11
Text to speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality…
Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis
Jian Cong, Shan Yang, Lei Xie +1
Current two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectr…
Data Efficient Voice Cloning from Noisy Samples with Domain Adversarial Training
Jian Cong, Shan Yang, Lei Xie +2
Data efficient voice cloning aims at synthesizing target speaker's voice with only a few enrollment samples at hand. To this end, speaker adaptation and speaker encoding are two ty…