1 paper
Marius Aasan, Odd Kolbjørnsen, Anne Schistad Solberg +1
Vision Transformer (ViT) architectures traditionally employ a grid-based approach to tokenization independent of the semantic content of an image. We propose a modular superpixel t…