1 paper
Zixin Wang, Dong Gong, Sen Wang +2
Contrastive Language-Image Pretraining (CLIP) excels at learning generalizable image representations but often falls short in zero-shot inference on certain downstream datasets. Te…