2 citations · 4 across the 5 of their papers we have counts for
1 paper · 1 filter
Lincan Cai, Jingxuan Kang, Shuang Li +4
Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmen…