3 papers
cs.CV2026
What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers
Moritz Pawlowsky, Antonis Vamvakeros, Alexander Weiss +3
Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as…
cs.CV2025
Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
Ronan Docherty, Antonis Vamvakeros, Samuel J. Cooper
Feature foundation models - usually vision transformers - offer rich semantic descriptors of images, useful for downstream tasks such as (interactive) segmentation and object detec…
cs.CV2025
Upsampling DINOv2 features for unsupervised vision tasks and weakly supervised materials segmentation
Ronan Docherty, Antonis Vamvakeros, Samuel J. Cooper
The features of self-supervised vision transformers (ViTs) contain strong semantic and positional information relevant to downstream tasks like object localization and segmentation…