activity
20182022
most citedA Survey on Neural Speech Synthesis

185 citations · 253 across the 16 of their papers we have counts for

collaborators

20 papers

cs.SD2022

A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

Haohan Guo, Fenglong Xie, Frank K. Soong +2

We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is us…

cs.SD2022

ParaTTS: Learning Linguistic and Prosodic Cross-sentence Information in Paragraph-based TTS

Liumeng Xue, Frank K. Soong, Shaofei Zhang +1

Recent advancements in neural end-to-end TTS models have shown high-quality, natural synthesized speech in a conventional sentence-based TTS. However, it is still challenging to re…

eess.AS202235 cited

NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality

Xu Tan, Jiawei Chen, Haohe Liu +11

Text to speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality…

cs.SD2022

Disentangling Style and Speaker Attributes for TTS Style Transfer

Xiaochun An, Frank K. Soong, Lei Xie

End-to-end neural TTS has shown improved performance in speech style transfer. However, the improvement is still limited by the available training data in both target styles and sp…

eess.AS2021185 cited

A Survey on Neural Speech Synthesis

Xu Tan, Tao Qin, Frank Soong +1

Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning co…

cs.SD20211 cited

Improving Performance of Seen and Unseen Speech Style Transfer in End-to-end Neural TTS

Xiaochun An, Frank K. Soong, Lei Xie

End-to-end neural TTS training has shown improved performance in speech style transfer. However, the improvement is still limited by the training data in both target styles and spe…