6 citations · 6 across the 1 of their papers we have counts for
1 paper
Yutaro Yamada, Yingtian Tang, Yoyo Zhang +1
Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-to-text retrieval. However, such performance does not…