CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
arXiv:2107.00652
Abstract
We present CSWin Transformer, an efficient and effective Transformer-based backbone for general-purpose vision tasks. A challenging issue in Transformer design is that global self-attention is very expensive to compute whereas local self-attention often limits the field of interactions of each token. To address this issue, we develop the Cross-Shaped Window self-attention mechanism for computing self-attention in the horizontal and vertical stripes in parallel that form a cross-shaped window, with each stripe obtained by splitting the input feature into stripes of equal width. We provide a mathematical analysis of the effect of the stripe width and vary the stripe width for different layers of the Transformer network which achieves strong modeling capability while limiting the computation cost. We also introduce Locally-enhanced Positional Encoding (LePE), which handles the local positional information better than existing encoding schemes. LePE naturally supports arbitrary input resolutions, and is thus especially effective and friendly for downstream tasks. Incorporated with these designs and a hierarchical structure, CSWin Transformer demonstrates competitive performance on common vision tasks. Specifically, it achieves 85.4\% Top-1 accuracy on ImageNet-1K without any extra training data or label, 53.9 box AP and 46.4 mask AP on the COCO detection task, and 52.2 mIOU on the ADE20K semantic segmentation task, surpassing previous state-of-the-art Swin Transformer backbone by +1.2, +2.0, +1.4, and +2.0 respectively under the similar FLOPs setting. By further pretraining on the larger dataset ImageNet-21K, we achieve 87.5% Top-1 accuracy on ImageNet-1K and high segmentation performance on ADE20K with 55.7 mIoU. The code and models are available at https://github.com/microsoft/CSWin-Transformer.
References in corpus (21)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Transformer in Transformer
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Random Erasing Data Augmentation
- Generating Long Sequences with Sparse Transformers
- Conditional Positional Encodings for Vision Transformers
- Axial Attention in Multidimensional Transformers
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Stand-Alone Self-Attention in Vision Models
- CvT: Introducing Convolutions to Vision Transformers
- Training Vision Transformers for Image Retrieval
- Pre-Trained Image Processing Transformer
- Rethinking Attention with Performers
- TransReID: Transformer-based Object Re-Identification
- Sparse Sinkhorn Attention
- End-to-End Video Instance Segmentation with Transformers
- Multiscale Vision Transformers
- Compressive Transformers for Long-Range Sequence Modelling
- Dual Path Networks
Cited by in corpus (16)
- A Survey on Visual Transformer
- Transformers in Vision: A Survey
- Swin Transformer V2: Scaling Up Capacity and Resolution
- Uformer: A General U-Shaped Transformer for Image Restoration
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
- SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks
- SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- On the Integration of Self-Attention and Convolution
- M2MRF: Many-to-Many Reassembly of Features for Tiny Lesion Segmentation in Fundus Images
- Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
- 2nd Place Solution to Google Landmark Recognition Competition 2021
- Exploiting Spatial-Temporal Semantic Consistency for Video Scene Parsing
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- Global Interaction Modelling in Vision Transformer via Super Tokens