1 citations · 1 across the 1 of their papers we have counts for
1 paper
Cangxiong Chen, Vinay P. Namboodiri, Julian Padget
CLIP is a widely used foundational vision-language model that is used for zero-shot image recognition and other image-text alignment tasks. We demonstrate that CLIP is vulnerable t…