1 paper · 1 filter
Ke Jin, Wankou Yang
The large-scale pretrained model CLIP, trained on 400 million image-text pairs, offers a promising paradigm for tackling vision tasks, albeit at the image level. Later works, such…