1 paper
Mengxiao Tian, Xinxiao Wu, Shuo Yang
Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved remarkable success in represent…