1 paper
Kenan Jiang, Xuehai He, Ruize Xu +1
Contrastive Language-Image Pretraining (CLIP) has demonstrated great zero-shot performance for matching images and text. However, it is still challenging to adapt vision-lanaguage…