1 citations · 2 across the 17 of their papers we have counts for
1 paper · 2 filters
Jiaming Zhang, Xingjun Ma, Xin Wang +4
With the rapid advancement of multimodal learning, pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capacities in bridging the gap between visual…