2 papers
cs.CV2026
JAFAR: Jack up Any Feature at Any Resolution
Paul Couairon, Loick Chambon, Louis Serrano +3
Foundation Vision Encoders have become essential for a wide range of dense vision tasks. However, their low-resolution spatial feature outputs necessitate feature upsampling to pro…
cs.CV2024
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
Spyros Gidaris, Andrei Bursuc, Oriane Simeoni +4
Self-supervised learning can be used for mitigating the greedy needs of Vision Transformer networks for very large fully-annotated datasets. Different classes of self-supervised le…