79 citations · 193 across the 17 of their papers we have counts for
21 papers
PromptTTS: Controllable Text-to-Speech with Text Descriptions
Zhifang Guo, Yichong Leng, Yihan Wu +2
Using a text description as prompt to guide the generation of text or images (e.g., GPT-3 or DALLE-2) has drawn wide attention recently. Beyond text and image generation, in this w…
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
Xu Tan, Jiawei Chen, Haohe Liu +11
Text to speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality…
AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios
Yihan Wu, Xu Tan, Bohan Li +5
Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new s…
InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
Zehua Chen, Xu Tan, Ke Wang +4
Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses…
A study on the efficacy of model pre-training in developing neural text-to-speech system
Guangyan Zhang, Yichong Leng, Daxin Tan +5
In the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers' data is a common approach. However, in terms of ultimately achie…
A Light-weight contextual spelling correction model for customizing transducer-based speech recognition systems
Xiaoqiang Wang, Yanqing Liu, Sheng Zhao +1
It's challenging to customize transducer-based automatic speech recognition (ASR) system with context information which is dynamic and unavailable during model training. In this wo…