1 paper · 1 filter
Bin Kang, Shaoguo Wen, Yang Fan +6
While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the structural mismatch between…