3 papers
cs.LG2026
BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks
Ivan SaboliÄ, Marin OrÅ¡iÄ, Josip Å ariÄ +1
Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that this paradigm is highly vulner…
cs.CV2026
EAST: Early Action Prediction Sampling Strategy with Token Masking
Iva SoviÄ, Ivan MartinoviÄ, Marin OrÅ¡iÄ
Early action prediction seeks to anticipate an action before it fully unfolds, but limited visual evidence makes this task especially challenging. We introduce EAST, a simple and e…
cs.CV2025
DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
Ivan MartinoviÄ, Josip Å ariÄ, Marin OrÅ¡iÄ +2
Pixel-level annotation is expensive and time-consuming. Semi-supervised segmentation methods address this challenge by learning models on few labeled images alongside a large corpu…