1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Jiapeng Wang, Chengyu Wang, Xiaodan Wang +2
Large-scale pre-trained text-image models with dual-encoder architectures (such as CLIP) are typically adopted for various vision-language applications, including text-image retrie…