1 paper
Zhixiang Chi, Yanan Wu, Li Gu +5
CLIP exhibits strong visual-textual alignment but struggle with open-vocabulary segmentation due to poor localization. Prior methods enhance spatial coherence by modifying intermed…