2 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Arsham Gholamzadeh Khoee, Yinan Yu, Robert Feldt
Large pre-trained vision-language models like CLIP have transformed computer vision by aligning images and text in a shared feature space, enabling robust zero-shot transfer via pr…