1 paper
Akang Wang, Xili Deng, Zhanxuan Hu +3
Vision-Language Models such as CLIP exhibit strong zero-shot recognition capability by aligning images with textual concepts, yet they often underperform on multi-label recognition…