1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Shaunak Halbe, Junjiao Tian, K J Joseph +4
Vision-language models (VLMs) like CLIP have been cherished for their ability to perform zero-shot visual recognition on open-vocabulary concepts. This is achieved by selecting the…