3 citations · 3 across the 1 of their papers we have counts for
1 paper
Jinhao Li, Haopeng Li, Sarah Erfani +3
It has recently been discovered that using a pre-trained vision-language model (VLM), e.g., CLIP, to align a whole query image with several finer text descriptions generated by a l…