1 paper
Guangke Chen, Yuhui Wang, Shouling Ji +2
Modern text-to-speech (TTS) systems, particularly those built on Large Audio-Language Models (LALMs), generate high-fidelity speech that faithfully reproduces input text and mimics…