AS-MLP: An Axial Shifted MLP Architecture for Vision
arXiv:2107.08391
Abstract
An Axial Shifted MLP architecture (AS-MLP) is proposed in this paper. Different from MLP-Mixer, where the global spatial feature is encoded for information flow through matrix transposition and one token-mixing MLP, we pay more attention to the local features interaction. By axially shifting channels of the feature map, AS-MLP is able to obtain the information flow from different axial directions, which captures the local dependencies. Such an operation enables us to utilize a pure MLP architecture to achieve the same local receptive field as CNN-like architecture. We can also design the receptive field size and dilation of blocks of AS-MLP, etc, in the same spirit of convolutional neural networks. With the proposed AS-MLP architecture, our model obtains 83.3% Top-1 accuracy with 88M parameters and 15.2 GFLOPs on the ImageNet-1K dataset. Such a simple yet effective architecture outperforms all MLP-based architectures and achieves competitive performance compared to the transformer-based architectures (e.g., Swin Transformer) even with slightly lower FLOPs. In addition, AS-MLP is also the first MLP-based architecture to be applied to the downstream tasks (e.g., object detection and semantic segmentation). The experimental results are also impressive. Our proposed AS-MLP obtains 51.5 mAP on the COCO validation set and 49.5 MS mIoU on the ADE20K dataset, which is competitive compared to the transformer-based architectures. Our AS-MLP establishes a strong baseline of MLP-based architecture. Code is available at https://github.com/svip-lab/AS-MLP.
Accepted by ICLR2022
References in corpus (15)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MLP-Mixer: An all-MLP Architecture for Vision
- Transformer in Transformer
- Conditional Positional Encodings for Vision Transformers
- DeepViT: Towards Deeper Vision Transformer
- LocalViT: Analyzing Locality in Vision Transformers
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet
- Long-Short Transformer: Efficient Transformers for Language and Vision
- Container: Context Aggregation Network
- Pay Attention to MLPs
- S-MLP: Spatial-Shift MLP Architecture for Vision
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
Cited by in corpus (12)
- Centralized Feature Pyramid for Object Detection
- MetaFormer Baselines for Vision
- CycleMLP: A MLP-like Architecture for Dense Prediction
- RepMLP: Re-parameterizing Convolutions into Fully-connected Layers for Image Recognition
- S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
- Adaptive Fourier Neural Operators: Efficient Token Mixers for Transformers
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- An Image Patch is a Wave: Phase-Aware Vision MLP
- RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
- PointMixer: MLP-Mixer for Point Cloud Understanding