2 papers
cs.SD2023
High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models
Chunyu Qiang, Hao Li, Yixin Tian +4
Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis deco…
eess.AS2023
Learning Speech Representation From Contrastive Token-Acoustic Pretraining
Chunyu Qiang, Hao Li, Yixin Tian +4
For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate…