4 papers
Suppressing Non-Semantic Noise in Masked Image Modeling Representations
Martine Hjelkrem-Tan, Marius Aasan, Rwiddhi Chakraborty +3
Masked Image Modeling (MIM) has become a ubiquitous self-supervised vision paradigm. In this work, we show that MIM objectives cause the learned representations to retain non-seman…
SPoT: Subpixel Placement of Tokens in Vision Transformers
Martine Hjelkrem-Tan, Marius Aasan, Gabriel Y. Arteaga +1
Vision Transformers naturally accommodate sparsity, yet standard tokenization methods confine features to discrete patch grids. This constraint prevents models from fully exploitin…
Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
Gabriel Y. Arteaga, Marius Aasan, Rwiddhi Chakraborty +4
Prototypical self-supervised learning methods consistently suffer from partial prototype collapse, where multiple prototypes converge to nearly identical representations. This unde…
Differentiable Hierarchical Visual Tokenization
Marius Aasan, Martine Hjelkrem-Tan, Nico Catalano +2
Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adap…