1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Mengjie Zhao, Junya Ono, Zhi Zhong +7
Contrastive cross-modal models such as CLIP and CLAP aid various vision-language (VL) and audio-language (AL) tasks. However, there has been limited investigation of and improvemen…