1 paper
Jiacheng Cheng, Hijung Valentina Shin, Nuno Vasconcelos +2
In the recent years, the dual-encoder vision-language models (\eg CLIP) have achieved remarkable text-to-image retrieval performance. However, we discover that these models usually…