2 papers
cs.SD2024
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
Sijing Chen, Yuan Feng, Laipeng He +22
With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin Audi…
cs.SD2022
Improving Cross-lingual Speech Synthesis with Triplet Training Scheme
Jianhao Ye, Hongbin Zhou, Zhiba Su +4
Recent advances in cross-lingual text-to-speech (TTS) made it possible to synthesize speech in a language foreign to a monolingual speaker. However, there is still a large gap betw…