1 paper
Khanh Binh Nguyen, Chae Jung Park
Large-scale pre-trained image-text models exhibit robust multimodal representations, yet applying the Contrastive Language-Image Pre-training (CLIP) model to audio-visual localizat…