9 citations · 9 across the 2 of their papers we have counts for
1 paper · 1 filter
Jiaming Lei, Lin Li, Chunping Wang +2
Benefiting from strong generalization ability, pre-trained vision language models (VLMs), e.g., CLIP, have been widely utilized in zero-shot scene understanding. Unlike simple reco…