2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.SD2025
Next Tokens Denoising for Speech Synthesis
Yanqing Liu, Ruiqing Xue, Chong Zhang +7
While diffusion and autoregressive (AR) models have significantly advanced generative modeling, they each present distinct limitations. AR models, which rely on causal attention, c…
eess.AS2023★ 2 cited
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model
Ruiqing Xue, Yanqing Liu, Lei He +4
Neural text-to-speech (TTS) generally consists of cascaded architecture with separately optimized acoustic model and vocoder, or end-to-end architecture with continuous mel-spectro…
cs.SD2022★ 1 cited
DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
Yanqing Liu, Ruiqing Xue, Lei He +2
Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from tw…