1 paper
Wenbo Zhang, Yifan Zhang, Jianfeng Lin +3
Pre-trained vision-language (V-L) models such as CLIP have shown excellent performance in many downstream cross-modal tasks. However, most of them are only applicable to the Englis…