4 papers
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
Tim Lebailly, Vijay Veerabadran, Satwik Kottur +2
Generative vision-language models (VLMs) exhibit strong high-level image understanding but lack spatially dense alignment between vision and language modalities, as our findings in…
Object-Centric Pretraining via Target Encoder Bootstrapping
Nikola Đukić, Tim Lebailly, Tinne Tuytelaars
Object-centric representation learning has recently been successfully applied to real-world datasets. This success can be attributed to pretrained non-object-centric foundation mod…
CrOC: Cross-View Online Clustering for Dense Visual Representation Learning
Thomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar +2
Learning dense visual representations without labels is an arduous task and more so from scene-centric data. We propose to tackle this challenging problem by proposing a Cross-view…
Global-Local Self-Distillation for Visual Representation Learning
Tim Lebailly, Tinne Tuytelaars
The downstream accuracy of self-supervised methods is tightly linked to the proxy task solved during training and the quality of the gradients extracted from it. Richer and more me…