1 paper · 1 filter
Tianrui Wang, Meng Ge, Cheng Gong +11
While LLM-based TTS models exhibit zero-shot emotion and speaker cloning, their cloning fidelity and pronunciation clarity degrade on unseen domains. Fine-tuning is essential for a…