1 paper
Yifei Ming, Yixuan Li
Pre-trained contrastive vision-language models have demonstrated remarkable performance across a wide range of tasks. However, they often struggle on fine-trained datasets with cat…