A survey of the Vision Transformers and their CNN-Transformer based Variants
arXiv:2305.09880 · doi:10.1007/s10462-023-10595-0
Abstract
Vision transformers have become popular as a possible substitute to convolutional neural networks (CNNs) for a variety of computer vision applications. These transformers, with their ability to focus on global relationships in images, offer large learning capacity. However, they may suffer from limited generalization as they do not tend to model local correlation in images. Recently, in vision transformers hybridization of both the convolution operation and self-attention mechanism has emerged, to exploit both the local and global image representations. These hybrid vision transformers, also referred to as CNN-Transformer architectures, have demonstrated remarkable results in vision applications. Given the rapidly growing number of hybrid vision transformers, it has become necessary to provide a taxonomy and explanation of these hybrid architectures. This survey presents a taxonomy of the recent vision transformer architectures and more specifically that of the hybrid vision transformers. Additionally, the key features of these architectures such as the attention mechanisms, positional embeddings, multi-scale processing, and convolution are also discussed. In contrast to the previous survey papers that are primarily focused on individual vision transformer architectures or CNNs, this survey uniquely emphasizes the emerging trend of hybrid vision transformers. By showcasing the potential of hybrid vision transformers to deliver exceptional performance across a range of computer vision tasks, this survey sheds light on the future directions of this rapidly evolving architecture.
Pages: 84, Figures: 16
References in corpus (26)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- Vision Transformers for Single Image Dehazing
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Facial Expression Recognition with Visual Transformers and Attentional Selective Fusion
- LocalViT: Analyzing Locality in Vision Transformers
- GasHis-Transformer: A Multi-scale Visual Transformer Approach for Gastric Histopathological Image Detection
- Adversarial Text-to-Image Synthesis: A Review
- CTCNet: A CNN-Transformer Cooperation Network for Face Image Super-Resolution
- CvT: Introducing Convolutions to Vision Transformers
- Training data-efficient image transformers & distillation through attention
- ResT: An Efficient Transformer for Visual Recognition
- DPT: Deformable Patch-based Transformer for Visual Recognition
- TransCAM: Transformer Attention-based CAM Refinement for Weakly Supervised Semantic Segmentation
- A Survey of Deep Learning Techniques for the Analysis of COVID-19 and their usability for Detecting Omicron
- Rethinking Spatial Dimensions of Vision Transformers
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation
- MaxViT: Multi-Axis Vision Transformer
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification
- Vision Transformers for Small Histological Datasets Learned through Knowledge Distillation
- EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications
- CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images
- HaloAE: An HaloNet based Local Transformer Auto-Encoder for Anomaly Detection and Localization
- BossNAS: Exploring Hybrid CNN-transformers with Block-wisely Self-supervised Neural Architecture Search
Cited by in corpus (7)
- EHCTNet: Enhanced Hybrid of CNN and Transformer Network for Remote Sensing Image Change Detection
- Predicting Mechanical Properties from Microstructure Images in Fiber-reinforced Polymers using Convolutional Neural Networks
- CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images
- Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring
- Universal Design Methodology for Printable Microstructural Materials via a New Deep Generative Learning Model: Application to a Piezocomposite
- Review of Fruit Tree Image Segmentation
- Local Concept Embeddings for Analysis of Concept Distributions in Vision DNN Feature Spaces