2 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Bikang Pan, Qun Li, Xiaoying Tang +6
The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite…