6 papers
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
Naomi Kombol, Ivan MartinoviÄ, SiniÅ¡a Å egviÄ +1
Foundational Vision Transformers (ViTs) have limited effectiveness in tasks requiring fine-grained spatial understanding, due to their fixed pre-training resolution and inherently…
Sequential keypoint density estimator: an overlooked baseline of skeleton-based video anomaly detection
Anja DeliÄ, Matej GrciÄ, SiniÅ¡a Å egviÄ
Detecting anomalous human behaviour is an important visual task in safety-critical applications such as healthcare monitoring, workplace safety, or public surveillance. In these co…
Seal Your Backdoor with Variational Defense
Ivan SaboliÄ, Matej GrciÄ, SiniÅ¡a Å egviÄ
We propose VIBE, a model-agnostic framework that trains classifiers resilient to backdoor attacks. The key concept behind our approach is to treat malicious inputs and corrupted la…
What Holds Back Open-Vocabulary Segmentation?
Josip Å ariÄ, Ivan MartinoviÄ, Matej Kristan +1
Standard segmentation setups are unable to deliver models that can recognize concepts outside the training taxonomy. Open-vocabulary approaches promise to close this gap through la…
DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
Ivan MartinoviÄ, Josip Å ariÄ, Marin OrÅ¡iÄ +2
Pixel-level annotation is expensive and time-consuming. Semi-supervised segmentation methods address this challenge by learning models on few labeled images alongside a large corpu…
A Survey on Training-free Open-Vocabulary Semantic Segmentation
Naomi Kombol, Ivan MartinoviÄ, SiniÅ¡a Å egviÄ
Semantic segmentation is one of the most fundamental tasks in image understanding with a long history of research, and subsequently a myriad of different approaches. Traditional me…