2 papers
cs.CV2026
Repurposing CLIP to Localize at Pixel Level
Jiaxiang Fang, Shiqiang Ma, Jing Wang +3
Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level de…
cs.CV2026
DualMem: Bypassing the Objectness Bottleneck for Calibrated Unknown-Stream Filtering in Open-World Object Detection
Yingjun Xiao, Xi Chen, Gang Fang +1
Open-world object detection (OWOD) requires detectors to localize known classes while identifying unknown objects for future incremental learning. We find that the unknown predicti…