1 paper · 1 filter
Jiawei Kong, Hao Fang, Sihang Guo +5
While pre-trained Vision-Language Models (VLMs) such as CLIP exhibit impressive representational capabilities for multimodal data, recent studies have revealed their vulnerability…