1 paper · 1 filter
Wanqing Cui, Rui Cheng, Jiafeng Guo +1
Existing two-stream models, such as CLIP, encode images and text through independent representations, showing good performance while ensuring retrieval speed, have attracted attent…