1 paper
Moritz Pawlowsky, Antonis Vamvakeros, Alexander Weiss +3
Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as…