21 citations · 28 across the 3 of their papers we have counts for
1 paper · 1 filter
Cheng-En Wu, Yu Tian, Haichao Yu +4
Vision-language models such as CLIP learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through…