1 paper
Yan Huang, Guowei Wang, Xu Wang +2
Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shif…