Global Interaction Modelling in Vision Transformer via Super Tokens
arXiv:2111.13156
Abstract
With the popularity of Transformer architectures in computer vision, the research focus has shifted towards developing computationally efficient designs. Window-based local attention is one of the major techniques being adopted in recent works. These methods begin with very small patch size and small embedding dimensions and then perform strided convolution (patch merging) in order to reduce the feature map size and increase embedding dimensions, hence, forming a pyramidal Convolutional Neural Network (CNN) like design. In this work, we investigate local and global information modelling in transformers by presenting a novel isotropic architecture that adopts local windows and special tokens, called Super tokens, for self-attention. Specifically, a single Super token is assigned to each image window which captures the rich local details for that window. These tokens are then employed for cross-window communication and global representation learning. Hence, most of the learning is independent of the image patches in the higher layers, and the class embedding is learned solely based on the Super tokens where is the window size. In standard image classification on Imagenet-1K, the proposed Super tokens based transformer (STT-S25) achieves 83.5\% accuracy which is equivalent to Swin transformer (Swin-B) with circa half the number of parameters (49M) and double the inference time throughput. The proposed Super token transformer offers a lightweight and promising backbone for visual recognition tasks.
References in corpus (14)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MLP-Mixer: An all-MLP Architecture for Vision
- Transformer in Transformer
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Reformer: The Efficient Transformer
- XCiT: Cross-Covariance Image Transformers
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- CycleMLP: A MLP-like Architecture for Dense Prediction
- RegionViT: Regional-to-Local Attention for Vision Transformers
- S-MLP: Spatial-Shift MLP Architecture for Vision
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- Rethinking Token-Mixing MLP for MLP-based Vision Backbone
- Hire-MLP: Vision MLP via Hierarchical Rearrangement