11 citations · 34 across the 11 of their papers we have counts for
1 paper · 1 filter
Bowen Shi, Peisen Zhao, Zichen Wang +8
Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual…