1 paper · 1 filter
Zhixiang Chi, Yanan Wu, Li Gu +5
CLIP exhibits strong visual-textual alignment but struggle with open-vocabulary segmentation due to poor localization. Prior methods enhance spatial coherence by modifying intermed…