A survey of the Vision Transformers and their CNN-Transformer based Variants
arXiv:2305.09880 · doi:10.1007/s10462-023-10595-0
Abstract
Vision transformers have become popular as a possible substitute to convolutional neural networks (CNNs) for a variety of computer vision applications. These transformers, with their ability to focus on global relationships in images, offer large learning capacity. However, they may suffer from limited generalization as they do not tend to model local correlation in images. Recently, in vision transformers hybridization of both the convolution operation and self-attention mechanism has emerged, to exploit both the local and global image representations. These hybrid vision transformers, also referred to as CNN-Transformer architectures, have demonstrated remarkable results in vision applications. Given the rapidly growing number of hybrid vision transformers, it has become necessary to provide a taxonomy and explanation of these hybrid architectures. This survey presents a taxonomy of the recent vision transformer architectures and more specifically that of the hybrid vision transformers. Additionally, the key features of these architectures such as the attention mechanisms, positional embeddings, multi-scale processing, and convolution are also discussed. In contrast to the previous survey papers that are primarily focused on individual vision transformer architectures or CNNs, this survey uniquely emphasizes the emerging trend of hybrid vision transformers. By showcasing the potential of hybrid vision transformers to deliver exceptional performance across a range of computer vision tasks, this survey sheds light on the future directions of this rapidly evolving architecture.
Pages: 84, Figures: 16
References in corpus (7)
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- CvT: Introducing Convolutions to Vision Transformers
- Training data-efficient image transformers & distillation through attention
- DPT: Deformable Patch-based Transformer for Visual Recognition
- Vision Transformers for Small Histological Datasets Learned through Knowledge Distillation
- CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images
- HaloAE: An HaloNet based Local Transformer Auto-Encoder for Anomaly Detection and Localization
Cited by in corpus (6)
- EHCTNet: Enhanced Hybrid of CNN and Transformer Network for Remote Sensing Image Change Detection
- CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images
- Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring
- Review of Fruit Tree Image Segmentation
- Universal Design Methodology for Printable Microstructural Materials via a New Deep Generative Learning Model: Application to a Piezocomposite
- Local Concept Embeddings for Analysis of Concept Distributions in Vision DNN Feature Spaces