1 paper
Sooyoung Park, Arda Senocak, Joon Son Chung
Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approa…