1 paper
HaeJun Yoo, Yongseop Shin, Insung Lee +2
Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-…