7 citations · 13 across the 5 of their papers we have counts for
1 paper · 1 filter
Kuanrong Liu, Siyuan Liang, Cheng Qian +2
As a general-purpose vision-language pretraining model, CLIP demonstrates strong generalization ability in image-text alignment tasks and has been widely adopted in downstream appl…