1 paper
Ziquan Liu, Zhewei Zhu, Xuyang Shi
Open-vocabulary semantic segmentation (OVSS) is fundamentally hampered by the coarse, image-level representations of CLIP, which lack precise pixel-level details. Existing training…