1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Xiangyu Gao, Yu Dai, Benliu Qiu +3
Owing to large-scale image-text contrastive training, pre-trained vision language model (VLM) like CLIP shows superior open-vocabulary recognition ability. Most existing open-vocab…