1 paper · 1 filter
Yongjin Yang, Jongwoo Ko, Se-Young Yun
Vision-language models (VLMs) like CLIP have demonstrated remarkable applicability across a variety of downstream tasks, including zero-shot image classification. Recently, the use…