1 paper
Pengfei Zhao, Rongbo Luan, Wei Zhang +2
Despite Contrastive Language-Image Pretraining (CLIP)'s remarkable capability to retrieve content across modalities, a substantial modality gap persists in its feature space. Intri…