1 paper
Beier Zhu, Yulei Niu, Yucheng Han +2
Thanks to the large pre-trained vision-language models (VLMs) like CLIP, we can craft a zero-shot classifier by "prompt", e.g., the confidence score of an image being "[CLASS]" can…