1 citations · 1 across the 1 of their papers we have counts for
1 paper
Junhui Yin, Xinyu Zhang, Lin Wu +1
Current pre-trained vision-language models, such as CLIP, have demonstrated remarkable zero-shot generalization capabilities across various downstream tasks. However, their perform…