4 papers · 1 filter
Suppressing Non-Semantic Noise in Masked Image Modeling Representations
Martine Hjelkrem-Tan, Marius Aasan, Rwiddhi Chakraborty +3
Masked Image Modeling (MIM) has become a ubiquitous self-supervised vision paradigm. In this work, we show that MIM objectives cause the learned representations to retain non-seman…
SPoT: Subpixel Placement of Tokens in Vision Transformers
Martine Hjelkrem-Tan, Marius Aasan, Gabriel Y. Arteaga +1
Vision Transformers naturally accommodate sparsity, yet standard tokenization methods confine features to discrete patch grids. This constraint prevents models from fully exploitin…
Differentiable Hierarchical Visual Tokenization
Marius Aasan, Martine Hjelkrem-Tan, Nico Catalano +2
Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adap…
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
Marius Aasan, Odd Kolbjørnsen, Anne Schistad Solberg +1
Vision Transformer (ViT) architectures traditionally employ a grid-based approach to tokenization independent of the semantic content of an image. We propose a modular superpixel t…