1 paper
Xuenan Xu, Xiaohang Xu, Zeyu Xie +3
Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events.…