1 citations · 2 across the 13 of their papers we have counts for
1 paper · 2 filters
Jiaming Zhang, Xin Wang, Xingjun Ma +3
Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data through joint embedding spaces.…