2 papers
cs.CV2026
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
Naomi Kombol, Ivan MartinoviÄ, SiniÅ¡a Å egviÄ +1
Foundational Vision Transformers (ViTs) have limited effectiveness in tasks requiring fine-grained spatial understanding, due to their fixed pre-training resolution and inherently…
cs.CV2025
A Survey on Training-free Open-Vocabulary Semantic Segmentation
Naomi Kombol, Ivan MartinoviÄ, SiniÅ¡a Å egviÄ
Semantic segmentation is one of the most fundamental tasks in image understanding with a long history of research, and subsequently a myriad of different approaches. Traditional me…