1 paper
Hongyu Hu, Tiancheng Lin, Jie Wang +2
Large-scale vision-language models (VLMs), e.g., CLIP, learn broad visual concepts from tedious training data, showing superb generalization ability. Amount of prompt learning meth…