3 papers
cs.CV2026
Repurposing CLIP to Localize at Pixel Level
Jiaxiang Fang, Shiqiang Ma, Jing Wang +3
Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level de…
cs.CV2026
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
Zhenyu Xie, Ji Xia, Michael Kampffmeyer +9
Controllable character animation has advanced rapidly in recent years, yet multi-character animation remains underexplored. As the number of characters grows, multi-character refer…
cs.CV2026
DCP-CLIP:A Coarse-to-Fine Framework for Open-Vocabulary Semantic Segmentation with Dual Interaction
Jing Wang, Huimin Shi, Quan Zhou +3
The recent years have witnessed the remarkable development for open-vocabulary semantic segmentation (OVSS) using visual-language foundation models, yet still suffer from following…