6 papers
Whareformer: Learning to Track What is Where in Long Egocentric Videos
Jacob Chalk, Saptarshi Sinha, Dima Damen +2
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining kno…
Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks
Pau de Jorge, César Roberto de Souza, Björn Michele +5
Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, m…
ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
Pavel Suma, Giorgos Kordopatis-Zilos, Yannis Kalantidis +1
Large-scale instance-level training data is scarce, so models are typically trained on domain-specific datasets. Yet in real-world retrieval, they must handle diverse domains, maki…
LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation
Vladan StojniÄ, Yannis Kalantidis, JiÅÃ Matas +1
We propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs). Our approach enhances the initial per-patch predictions of VLMs…
DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
Mert Bulent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas +3
Recent multi-teacher distillation methods have unified the encoders of multiple foundation models into a single encoder, achieving competitive performance on core vision tasks like…
UNIC: Universal Classification Models via Multi-teacher Distillation
Mert Bulent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas +2
Pretrained models have become a commodity and offer strong results on a broad range of tasks. In this work, we focus on classification and seek to learn a unique encoder able to ta…