1 paper · 1 filter
Juan Yeo, Soonwoo Cha, Jiwoo Song +2
Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still…