1 paper
Samyak Rawlekar, Amitabh Swain, Yujun Cai +3
Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in [CLS] token attention maps of the final layer. However, the…