Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
arXiv:2402.02941 · doi:10.1016/j.engappai.2025.110057
Abstract
The hybrid of Convolutional Neural Network (CNN) and Vision Transformers (ViT) architectures has emerged as a groundbreaking approach, pushing the boundaries of computer vision (CV). This comprehensive review provides a thorough examination of the literature on state-of-the-art hybrid CNN-ViT architectures, exploring the synergies between these two approaches. The main content of this survey includes: (1) a background on the vanilla CNN and ViT, (2) systematic review of various taxonomic hybrid designs to explore the synergy achieved through merging CNNs and ViTs models, (3) comparative analysis and application task-specific synergy between different hybrid architectures, (4) challenges and future directions for hybrid models, (5) lastly, the survey concludes with a summary of key findings and recommendations. Through this exploration of hybrid CV architectures, the survey aims to serve as a guiding resource, fostering a deeper understanding of the intricate dynamics between CNNs and ViTs and their collective impact on shaping the future of CV architectures.
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Linformer: Self-Attention with Linear Complexity
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- EfficientFormer: Vision Transformers at MobileNet Speed
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- A comparative study between vision transformers and CNNs in digital pathology
- Dilated Neighborhood Attention Transformer
- Attention Modules Improve Image-Level Anomaly Detection for Industrial Inspection: A DifferNet Case Study