3 papers
cs.CV2026
Suppressing Non-Semantic Noise in Masked Image Modeling Representations
Martine Hjelkrem-Tan, Marius Aasan, Rwiddhi Chakraborty +3
Masked Image Modeling (MIM) has become a ubiquitous self-supervised vision paradigm. In this work, we show that MIM objectives cause the learned representations to retain non-seman…
cs.CV2026
SPoT: Subpixel Placement of Tokens in Vision Transformers
Martine Hjelkrem-Tan, Marius Aasan, Gabriel Y. Arteaga +1
Vision Transformers naturally accommodate sparsity, yet standard tokenization methods confine features to discrete patch grids. This constraint prevents models from fully exploitin…
cs.CV2025
Differentiable Hierarchical Visual Tokenization
Marius Aasan, Martine Hjelkrem-Tan, Nico Catalano +2
Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adap…