1 paper
Jeannie Chung, Hanna Jang, Ingyeong Yang +2
CLIP aligns image and text embeddings via contrastive learning and demonstrates strong zero-shot generalization. Its large-scale architecture requires substantial computational and…