7 papers
Beyond Natural-Image Foundation Models: Benchmarking Satellite Pretraining for Ophthalmic Image Analysis
Lovre Antonio Budimir, Mingya Alexa Gong, Alyssa Foong Quinney +5
Vision Foundation Models (VFMs) have emerged as a promising approach in medical imaging, producing broadly applicable systems that can be efficiently adapted across diverse imaging…
Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging
Mingya Alexa Gong, Da Ma, Lovre Antonio Budimir +7
Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence…
BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks
Ivan SaboliÄ, Marin OrÅ¡iÄ, Josip Å ariÄ +1
Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that this paradigm is highly vulner…
Sparse Code Uplifting for Efficient 3D Language Gaussian Splatting
Lovre Antonio Budimir, Yushi Guan, Steve Ryhner +2
3D Language Gaussian Splatting (3DLGS) augments 3D Gaussian Splatting with language-aligned visual features for open-vocabulary 3D scene understanding. A core challenge is efficien…
OOS-DSD: Improving Out-of-stock Detection in Retail Images using Auxiliary Tasks
Franko Å ikiÄ, Sven LonÄariÄ
Out-of-stock (OOS) detection is a very important retail verification process that aims to infer the unavailability of products in their designated areas on the shelf. In this paper…
DHECA-SuperGaze: Dual Head-Eye Cross-Attention and Super-Resolution for Unconstrained Gaze Estimation
Franko Å ikiÄ, Donik VrÅ¡nak, Sven LonÄariÄ
Unconstrained gaze estimation is the process of determining where a subject is directing their visual attention in uncontrolled environments. Gaze estimation systems are important…