5 papers · 1 filter
SA-VIS: Sparse frame Annotations for training Video Instance Segmentation
Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain +2
Recent online video instance segmentation (VIS) methods have achieved impressive results, thus becoming the preferred approach to segment instances in videos. Despite the resurgenc…
VoxCor: Training-Free Volumetric Features for Multimodal Voxel Correspondence
Guney Tombak, Ertunc Erdil, Ender Konukoglu
Cross-modal 3D medical image analysis requires voxelwise representations that remain anatomically consistent across imaging contrasts, scanners, and acquisition protocols. Recent w…
Trustworthy Endoscopic Super-Resolution
Julio Silva-RodrÃguez, Ender Konukoglu
Super-resolution (SR) models are attracting growing interest for enhancing minimally invasive surgery and diagnostic videos under hardware constraints. However, valid concerns rema…
Spatial Autoregressive Modeling of DINOv3 Embeddings for Unsupervised Anomaly Detection
Ertunc Erdil, Nico Schulthess, Guney Tombak +1
DINO models provide rich patch-level representations that have recently enabled strong performance in unsupervised anomaly detection (UAD). Most existing methods extract patch embe…
Semi-Supervised Few-Shot Adaptation of Vision-Language Models
Julio Silva-RodrÃguez, Ender Konukoglu
Vision-language models (VLMs) pre-trained on large, heterogeneous data sources are becoming increasingly popular, providing rich multi-modal embeddings that enable efficient transf…