1 paper
Cheng Cheng, Lin Song, Ruoyi Xue +4
The contrastive vision-language pre-training, known as CLIP, demonstrates remarkable potential in perceiving open-world visual concepts, enabling effective zero-shot image recognit…