1 citations · 1 across the 1 of their papers we have counts for
1 paper
Zhifeng Kong, Sang-gil Lee, Deepanway Ghosal +5
It is an open challenge to obtain high quality training data, especially captions, for text-to-audio models. Although prior methods have leveraged \textit{text-only language models…