LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
arXiv:2104.01136
Abstract
We design a family of image classification architectures that optimize the trade-off between accuracy and efficiency in a high-speed regime. Our work exploits recent findings in attention-based architectures, which are competitive on highly parallel processing hardware. We revisit principles from the extensive literature on convolutional neural networks to apply them to transformers, in particular activation maps with decreasing resolutions. We also introduce the attention bias, a new way to integrate positional information in vision transformers. As a result, we propose LeVIT: a hybrid neural network for fast inference image classification. We consider different measures of efficiency on different hardware platforms, so as to best reflect a wide range of application scenarios. Our extensive experiments empirically validate our technical choices and show they are suitable to most architectures. Overall, LeViT significantly outperforms existing convnets and vision transformers with respect to the speed/accuracy tradeoff. For example, at 80% ImageNet top-1 accuracy, LeViT is 5 times faster than EfficientNet on CPU. We release the code at https://github.com/facebookresearch/LeViT
References in corpus (15)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- EfficientNetV2: Smaller Models and Faster Training
- Generating Long Sequences with Sparse Transformers
- Do ImageNet Classifiers Generalize to ImageNet?
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- High-Performance Large-Scale Image Recognition Without Normalization
- CvT: Introducing Convolutions to Vision Transformers
- Toward Transformer-Based Object Detection
- Training Vision Transformers for Image Retrieval
- Are we done with ImageNet?
- Rethinking Spatial Dimensions of Vision Transformers
- Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model
- MultiGrain: a unified image embedding for classes and instances
- Global Self-Attention Networks for Image Recognition
Cited by in corpus (28)
- A Survey on Visual Transformer
- Transformers in Vision: A Survey
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- XCiT: Cross-Covariance Image Transformers
- MISSFormer: An Effective Medical Image Segmentation Transformer
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- RegionViT: Regional-to-Local Attention for Vision Transformers
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition
- Vision Xformers: Efficient Attention for Image Classification
- Searching for Efficient Multi-Stage Vision Transformers
- FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- Shunted Self-Attention via Multi-Scale Token Aggregation
- Towards Robust Vision Transformer
- CAPE: Encoding Relative Positions with Continuous Augmented Positional Embeddings
- Styleformer: Transformer based Generative Adversarial Networks with Style Vector
- KVT: k-NN Attention for Boosting Vision Transformers
- Transformed CNNs: recasting pre-trained convolutional layers with self-attention
- Armour: Generalizable Compact Self-Attention for Vision Transformers
- Exploring and Improving Mobile Level Vision Transformers