MetaFormer Baselines for Vision
arXiv:2210.13452 · doi:10.1109/TPAMI.2023.3329173
Abstract
MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capacity of MetaFormer, again, without focusing on token mixer design: we introduce several baseline models under MetaFormer using the most basic or common mixers, and summarize our observations as follows: (1) MetaFormer ensures solid lower bound of performance. By merely adopting identity mapping as the token mixer, the MetaFormer model, termed IdentityFormer, achieves >80% accuracy on ImageNet-1K. (2) MetaFormer works well with arbitrary token mixers. When specifying the token mixer as even a random matrix to mix tokens, the resulting model RandFormer yields an accuracy of >81%, outperforming IdentityFormer. Rest assured of MetaFormer's results when new token mixers are adopted. (3) MetaFormer effortlessly offers state-of-the-art results. With just conventional token mixers dated back five years ago, the models instantiated from MetaFormer already beat state of the art. (a) ConvFormer outperforms ConvNeXt. Taking the common depthwise separable convolutions as the token mixer, the model termed ConvFormer, which can be regarded as pure CNNs, outperforms the strong CNN model ConvNeXt. (b) CAFormer sets new record on ImageNet-1K. By simply applying depthwise separable convolutions as token mixer in the bottom stages and vanilla self-attention in the top stages, the resulting model CAFormer sets a new record on ImageNet-1K: it achieves an accuracy of 85.5% at 224x224 resolution, under normal supervised training without external data or distillation. In our expedition to probe MetaFormer, we also find that a new activation, StarReLU, reduces 71% FLOPs of activation compared with GELU yet achieves better performance. We expect StarReLU to find great potential in MetaFormer-like models alongside other neural networks.
Accepted to TPAMI. Code: https://github.com/sail-sg/metaformer
References in corpus (17)
- Adam: A Method for Stochastic Optimization
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Decoupled Weight Decay Regularization
- Gaussian Error Linear Units (GELUs)
- Longformer: The Long-Document Transformer
- PaLM: Scaling Language Modeling with Pathways
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- TENER: Adapting Transformer Encoder for Named Entity Recognition
- Vision GNN: An Image is Worth Graph of Nodes
- HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions
- Focal Modulation Networks
- UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- Sequencer: Deep LSTM for Image Classification
- Refiner: Refining Self-attention for Vision Transformers
- NormFormer: Improved Transformer Pretraining with Extra Normalization
Cited by in corpus (9)
- Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training
- FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
- Scaling up self-supervised learning for improved surgical foundation models
- HDBFormer: Efficient RGB-D Semantic Segmentation with A Heterogeneous Dual-Branch Framework
- Exploring the Effect of Dataset Diversity in Self-Supervised Learning for Surgical Computer Vision
- Multi-Modal Landslide Detection from Sentinel-1 SAR and Sentinel-2 Optical Imagery Using Multi-Encoder Vision Transformers and Ensemble Learning
- Dilated Convolution with Learnable Spacings: beyond bilinear interpolation
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- Accuracy Improvement of Cell Image Segmentation Using Feedback Former