1 paper
Jiawei Kong, Hao Fang, Sihang Guo +5
While pre-trained Vision-Language Models (VLMs) such as CLIP exhibit impressive representational capabilities for multimodal data, recent studies have revealed their vulnerability…