ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases
arXiv:2103.10697 · doi:10.1088/1742-5468/ac9830
Abstract
Convolutional architectures have proven extremely successful for vision tasks. Their hard inductive biases enable sample-efficient learning, but come at the cost of a potentially lower performance ceiling. Vision Transformers (ViTs) rely on more flexible self-attention layers, and have recently outperformed CNNs for image classification. However, they require costly pre-training on large external datasets or distillation from pre-trained convolutional networks. In this paper, we ask the following question: is it possible to combine the strengths of these two architectures while avoiding their respective limitations? To this end, we introduce gated positional self-attention (GPSA), a form of positional self-attention which can be equipped with a ``soft" convolutional inductive bias. We initialise the GPSA layers to mimic the locality of convolutional layers, then give each attention head the freedom to escape locality by adjusting a gating parameter regulating the attention paid to position versus content information. The resulting convolutional-like ViT architecture, ConViT, outperforms the DeiT on ImageNet, while offering a much improved sample efficiency. We further investigate the role of locality in learning by first quantifying how it is encouraged in vanilla self-attention layers, then analysing how it is escaped in GPSA layers. We conclude by presenting various ablations to better understand the success of the ConViT. Our code and models are released publicly at https://github.com/facebookresearch/convit.
References in corpus (3)
Cited by in corpus (35)
- Early Convolutions Help Transformers See Better
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- Class-Incremental Learning: A Survey
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Vision Transformer with Attentive Pooling for Robust Facial Expression Recognition
- How Deep Learning Sees the World: A Survey on Adversarial Attacks & Defenses
- Memorizing Complementation Network for Few-Shot Class-Incremental Learning
- SOFT: Softmax-free Transformer with Linear Complexity
- Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
- SCCAM: Supervised Contrastive Convolutional Attention Mechanism for Ante-hoc Interpretable Fault Diagnosis with Limited Fault Samples
- PHNNs: Lightweight Neural Networks via Parameterized Hypercomplex Convolutions
- Concurrent ischemic lesion age estimation and segmentation of CT brain using a Transformer-based network
- A Convolutional Vision Transformer for Semantic Segmentation of Side-Scan Sonar Data
- DnSwin: Toward Real-World Denoising via Continuous Wavelet Sliding-Transformer
- Medical Image Segmentation using LeViT-UNet++: A Case Study on GI Tract Data
- DiRecNetV2: A Transformer-Enhanced Network for Aerial Disaster Recognition
- Vision Transformers: From Semantic Segmentation to Dense Prediction
- DMTNet: Dynamic Multi-scale Network for Dual-pixel Images Defocus Deblurring with Transformer
- Pruning Self-attentions into Convolutional Layers in Single Path
- Softmax-free Linear Transformers
- PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
- BabyNet: Residual Transformer Module for Birth Weight Prediction on Fetal Ultrasound Video
- Parameterization of Cross-Token Relations with Relative Positional Encoding for Vision MLP
- Improving Robustness and Reliability in Medical Image Classification with Latent-Guided Diffusion and Nested-Ensembles
- gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
- Study of positional encoding approaches for Audio Spectrogram Transformers
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- Learning Broken Symmetries with Approximate Invariance
- POST: Photonic Swin Transformer for Automated and Efficient Prediction of PCSEL
- Token Pooling in Vision Transformers
- Weighted Ensemble Models Are Strong Continual Learners
- AlphaViT: A flexible game-playing AI for multiple games and variable board sizes
- Relaxed syntax modeling in Transformers for future-proof license plate recognition