19 citations · 20 across the 4 of their papers we have counts for
1 paper · 1 filter
Wenbo Zhang, Yifan Zhang, Jianfeng Lin +3
Pre-trained vision-language (V-L) models such as CLIP have shown excellent performance in many downstream cross-modal tasks. However, most of them are only applicable to the Englis…