1 paper
Yuxiao Chen, Jianbo Yuan, Yu Tian +5
Contrastive learning-based vision-language pre-training approaches, such as CLIP, have demonstrated great success in many vision-language tasks. These methods achieve cross-modal a…