Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
arXiv:2503.02891 · doi:10.1016/j.neucom.2025.130417
Abstract
In recent years, vision transformers (ViTs) have emerged as powerful and promising techniques for computer vision tasks such as image classification, object detection, and segmentation. Unlike convolutional neural networks (CNNs), which rely on hierarchical feature extraction, ViTs treat images as sequences of patches and leverage self-attention mechanisms. However, their high computational complexity and memory demands pose significant challenges for deployment on resource-constrained edge devices. To address these limitations, extensive research has focused on model compression techniques and hardware-aware acceleration strategies. Nonetheless, a comprehensive review that systematically categorizes these techniques and their trade-offs in accuracy, efficiency, and hardware adaptability for edge deployment remains lacking. This survey bridges this gap by providing a structured analysis of model compression techniques, software tools for inference on edge, and hardware acceleration strategies for ViTs. We discuss their impact on accuracy, efficiency, and hardware adaptability, highlighting key challenges and emerging research directions to advance ViT deployment on edge platforms, including graphics processing units (GPUs), application-specific integrated circuit (ASICs), and field-programmable gate arrays (FPGAs). The goal is to inspire further research with a contemporary guide on optimizing ViTs for efficient deployment on edge devices.
Accepted in Neurocomputing, Elsevier
References in corpus (20)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Knowledge Distillation: A Survey
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Towards Accurate Post-Training Quantization for Vision Transformer
- Vision Transformer Pruning
- Patch Similarity Aware Data-Free Quantization for Vision Transformers
- EasyQuant: Post-training Quantization via Scale Optimization
- MViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design
- VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer
- Learning Efficient Vision Transformers via Fine-Grained Manifold Distillation
- CP-ViT: Cascade Vision Transformer Pruning via Progressive Sparsity Prediction
- Auto-ViT-Acc: An FPGA-Aware Automatic Acceleration Framework for Vision Transformer with Mixed-Scheme Quantization
- PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
- LRP-QViT: Mixed-Precision Vision Transformer Quantization via Layer-wise Relevance Propagation
- Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
- Improving Post-Training Quantization on Object Detection with Task Loss-Guided Lp Metric
- Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference