1 paper
Haina Zhu, Yao Xiao, Xiquan Li +5
We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over at…