gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
arXiv:2208.11718 · doi:10.1109/ICASSP49357.2023.10096453
Abstract
Following the success in language domain, the self-attention mechanism (transformer) is adopted in the vision domain and achieving great success recently. Additionally, as another stream, multi-layer perceptron (MLP) is also explored in the vision domain. These architectures, other than traditional CNNs, have been attracting attention recently, and many methods have been proposed. As one that combines parameter efficiency and performance with locality and hierarchy in image recognition, we propose gSwin, which merges the two streams; Swin Transformer and (multi-head) gMLP. We showed that our gSwin can achieve better accuracy on three vision tasks, image classification, object detection and semantic segmentation, than Swin Transformer, with smaller model size.
5 pages, 7 figures, IEEE ICASSP 2023
References in corpus (9)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MLP-Mixer: An all-MLP Architecture for Vision
- Transformer in Transformer
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases
- Cascade R-CNN: Delving into High Quality Object Detection
- Visual Transformers: Token-based Image Representation and Processing for Computer Vision
- LocalViT: Analyzing Locality in Vision Transformers
- CycleMLP: A MLP-like Architecture for Dense Prediction
- AS-MLP: An Axial Shifted MLP Architecture for Vision