1 citations · 1 across the 1 of their papers we have counts for
1 paper
Cheng-En Wu, Yu Tian, Haichao Yu +4
Vision-language models such as CLIP learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through…