1 paper · 1 filter
Wenxuan Bao, Ruxi Deng, Jingrui He
Pretrained vision-language models such as CLIP achieve strong zero-shot generalization but remain vulnerable to distribution shifts caused by input corruptions. In this work, we in…