CoAtNet: Marrying Convolution and Attention for All Data Sizes
arXiv:2106.04803
Abstract
Transformers have attracted increasing interests in computer vision, but they still fall behind state-of-the-art convolutional networks. In this work, we show that while Transformers tend to have larger model capacity, their generalization can be worse than convolutional networks due to the lack of the right inductive bias. To effectively combine the strengths from both architectures, we present CoAtNets(pronounced "coat" nets), a family of hybrid models built from two key insights: (1) depthwise Convolution and self-Attention can be naturally unified via simple relative attention; (2) vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency. Experiments show that our CoAtNets achieve state-of-the-art performance under different resource constraints across various datasets: Without extra data, CoAtNet achieves 86.0% ImageNet top-1 accuracy; When pre-trained with 13M images from ImageNet-21K, our CoAtNet achieves 88.56% top-1 accuracy, matching ViT-huge pre-trained with 300M images from JFT-300M while using 23x less data; Notably, when we further scale up CoAtNet with JFT-3B, it achieves 90.88% top-1 accuracy on ImageNet, establishing a new state-of-the-art result.
References in corpus (6)
- Language Models are Few-Shot Learners
- DeepViT: Towards Deeper Vision Transformer
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- High-Performance Large-Scale Image Recognition Without Normalization
- Stand-Alone Self-Attention in Vision Models
- CvT: Introducing Convolutions to Vision Transformers
Cited by in corpus (10)
- HRFormer: High-Resolution Transformer for Dense Prediction
- A Practical Deep Learning-Based Acoustic Side Channel Attack on Keyboards
- PSLT: A Light-weight Vision Transformer with Ladder Self-Attention and Progressive Shift
- Synthetic Data Supervised Salient Object Detection
- Socially Enhanced Situation Awareness from Microblogs using Artificial Intelligence: A Survey
- Frequency Disentangled Features in Neural Image Compression
- Adversarial Token Attacks on Vision Transformers
- Parameterization of Cross-Token Relations with Relative Positional Encoding for Vision MLP
- Detecting Unknown DGAs without Context Information
- COVID-19 Pneumonia Severity Prediction using Hybrid Convolution-Attention Neural Architectures