1 paper
Chengyin Hu, Xiang Chen, Zhe Jia +4
Vision-Language Models (VLMs) are trained on image-text pairs collected under canonical visual conditions and achieve strong performance on multimodal tasks. However, their robustn…