CycleMLP: A MLP-like Architecture for Dense Prediction
arXiv:2107.10224
Abstract
This paper presents a simple MLP-like architecture, CycleMLP, which is a versatile backbone for visual recognition and dense predictions. As compared to modern MLP architectures, e.g., MLP-Mixer, ResMLP, and gMLP, whose architectures are correlated to image size and thus are infeasible in object detection and segmentation, CycleMLP has two advantages compared to modern approaches. (1) It can cope with various image sizes. (2) It achieves linear computational complexity to image size by using local windows. In contrast, previous MLPs have computations due to fully spatial connections. We build a family of models which surpass existing MLPs and even state-of-the-art Transformer-based models, e.g., Swin Transformer, while using fewer parameters and FLOPs. We expand the MLP-like models' applicability, making them a versatile backbone for dense prediction tasks. CycleMLP achieves competitive results on object detection, instance segmentation, and semantic segmentation. In particular, CycleMLP-Tiny outperforms Swin-Tiny by 1.3% mIoU on ADE20K dataset with fewer FLOPs. Moreover, CycleMLP also shows excellent zero-shot robustness on ImageNet-C dataset. Code is available at https://github.com/ShoufaChen/CycleMLP.
ICLR 2022 (Oral). Camera-ready Code: https://github.com/ShoufaChen/CycleMLP
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- Is Space-Time Attention All You Need for Video Understanding?
- MLP-Mixer: An all-MLP Architecture for Vision
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- CvT: Introducing Convolutions to Vision Transformers
- MDMMT: Multidomain Multimodal Transformer for Video Retrieval
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- Multiscale Vision Transformers
- Pay Attention to MLPs
- S-MLP: Spatial-Shift MLP Architecture for Vision
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- Global Filter Networks for Image Classification
Cited by in corpus (15)
- TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting
- Deep Instance Segmentation with Automotive Radar Detection Points
- S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
- MetaFormer Is Actually What You Need for Vision
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- Image and Model Transformation with Secret Key for Vision Transformer
- UniNeXt: Exploring A Unified Architecture for Vision Recognition
- PRSeg: A Lightweight Patch Rotate MLP Decoder for Semantic Segmentation
- Parameterization of Cross-Token Relations with Relative Positional Encoding for Vision MLP
- gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- An Image Patch is a Wave: Phase-Aware Vision MLP
- PointMixer: MLP-Mixer for Point Cloud Understanding
- RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
- Global Interaction Modelling in Vision Transformer via Super Tokens