38 citations · 78 across the 11 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2023★ 38 cited
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Kai Shen, Zeqian Ju, Xu Tan +6
Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, an…
eess.AS2023★ 2 cited
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model
Ruiqing Xue, Yanqing Liu, Lei He +4
Neural text-to-speech (TTS) generally consists of cascaded architecture with separately optimized acoustic model and vocoder, or end-to-end architecture with continuous mel-spectro…