1 citations · 1 across the 1 of their papers we have counts for
1 paper
Heeseong Shin, Chaehyun Kim, Sunghwan Hong +4
Large-scale vision-language models like CLIP have demonstrated impressive open-vocabulary capabilities for image-level tasks, excelling in recognizing what objects are present. How…