1 paper · 1 filter
Zhiyu Zhu, Zhibo Jin, Jiayu Zhang +4
The task of identifying multimodal image-text representations has garnered increasing attention, particularly with models such as CLIP (Contrastive Language-Image Pretraining), whi…