4 papers · 1 filter
DINOv3
Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23
Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. B…
Cluster and Predict Latent Patches for Improved Masked Image Modeling
Timothée Darcet, Federico Baldassarre, Maxime Oquab +2
Masked Image Modeling (MIM) offers a promising approach to self-supervised representation learning, however existing MIM models still lag behind the state-of-the-art. In this paper…
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
Cijo Jose, Théo Moutakanni, Dahyun Kang +11
Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models…
Vision Transformers Need Registers
Timothée Darcet, Maxime Oquab, Julien Mairal +1
Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised…