3 citations · 5 across the 9 of their papers we have counts for
1 paper · 1 filter
Mohan Li, Rama Doddipatla, Philip C. Woodland
Contrastive Language-Audio Pretraining (CLAP) aligns text and audio in a shared embedding space, but encoding each modality independently limits its ability to model cross-modal se…