1 paper · 1 filter
Arghya Pal, Sailaja Rajanala
Text-to-audio retrieval has made significant progress with shared embedding models such as CLAP and Pengi, yet they often struggle with fine-grained semantic alignment due to the i…