2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Ruyang Liu, Jingjia Huang, Ge Li +3
Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention f…