1 paper
KiHyun Nam, Jongmin Choi, Hyeongkeun Lee +2
Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with larg…