A Survey of Model Compression and Acceleration for Deep Neural Networks
arXiv:1710.09282
Abstract
Deep neural networks (DNNs) have recently achieved great success in many visual recognition tasks. However, existing deep neural network models are computationally expensive and memory intensive, hindering their deployment in devices with low memory resources or in applications with strict latency requirements. Therefore, a natural thought is to perform model compression and acceleration in deep networks without significantly decreasing the model performance. During the past five years, tremendous progress has been made in this area. In this paper, we review the recent techniques for compacting and accelerating DNN models. In general, these techniques are divided into four categories: parameter pruning and quantization, low-rank factorization, transferred/compact convolutional filters, and knowledge distillation. Methods of parameter pruning and quantization are described first, after that the other techniques are introduced. For each category, we also provide insightful analysis about the performance, related applications, advantages, and drawbacks. Then we go through some very recent successful methods, for example, dynamic capacity networks and stochastic depths networks. After that, we survey the evaluation matrices, the main datasets used for evaluating the model performance, and recent benchmark efforts. Finally, we conclude this paper, discuss remaining the challenges and possible directions for future work.
Published in IEEE Signal Processing Magazine, updated version including more recent works
References in corpus (19)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Neural Architecture Search with Reinforcement Learning
- Striving for Simplicity: The All Convolutional Net
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- FitNets: Hints for Thin Deep Nets
- DARTS: Differentiable Architecture Search
- Deep Learning with Limited Numerical Precision
- Compressing Deep Convolutional Networks using Vector Quantization
- Rethinking the Value of Network Pruning
- Trained Ternary Quantization
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Structured Sparsity in Deep Neural Networks
- Channel Pruning for Accelerating Very Deep Neural Networks
- ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Deep Pyramidal Residual Networks with Separated Stochastic Depth
- S3Pool: Pooling with Stochastic Spatial Sampling
Cited by in corpus (206)
- Res2Net: A New Multi-scale Backbone Architecture
- Pre-trained Models for Natural Language Processing: A Survey
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- Binary Neural Networks: A Survey
- RGB-D Salient Object Detection: A Survey
- Modality specific U-Net variants for biomedical image segmentation: A survey
- Layer-wise training convolutional neural networks with smaller filters for human activity recognition using wearable sensors
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Recurrent Neural Networks: An Embedded Computing Perspective
- Deep Learning in Mobile and Wireless Networking: A Survey
- Knowledge Distillation in Deep Learning and its Applications
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- Training for Faster Adversarial Robustness Verification via Inducing ReLU Stability
- Review of data analysis in vision inspection of power lines with an in-depth discussion of deep learning technology
- Deep -Means: Re-Training and Parameter Sharing with Harder Cluster Assignments for Compressing Deep Convolutions
- Dual Dynamic Inference: Enabling More Efficient, Adaptive and Controllable Deep Inference
- CheXtransfer: Performance and Parameter Efficiency of ImageNet Models for Chest X-Ray Interpretation
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Artificial Neural Networks for Photonic Applications: From Algorithms to Implementation
- Heterogeneous Multilayer Generalized Operational Perceptron
- Deep Learning Techniques for In-Crop Weed Identification: A Review
- Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
- Learned discretizations for passive scalar advection in a 2-D turbulent flow
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Integrated Photonic Tensor Processing Unit for a Matrix Multiply: a Review
- SLSNet: Skin lesion segmentation using a lightweight generative adversarial network
- Multi defect detection and analysis of electron microscopy images with deep learning
- Compressing GANs using Knowledge Distillation
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Unbiased Knowledge Distillation for Recommendation
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- Diet Code Is Healthy: Simplifying Programs for Pre-trained Models of Code
- Knowledge Distillation approach towards Melanoma Detection
- Resource-Efficient Deep Learning: A Survey on Model-, Arithmetic-, and Implementation-Level Techniques
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Focal Plane Wavefront Sensing using Machine Learning: Performance of Convolutional Neural Networks compared to Fundamental Limits
- SCSP: Spectral Clustering Filter Pruning with Soft Self-adaption Manners
- FPUS23: An Ultrasound Fetus Phantom Dataset with Deep Neural Network Evaluations for Fetus Orientations, Fetal Planes, and Anatomical Features
- Feature Fusion for Online Mutual Knowledge Distillation
- FedVision: An Online Visual Object Detection Platform Powered by Federated Learning
- Spiking neural networks trained with backpropagation for low power neuromorphic implementation of voice activity detection
- A Unified Lottery Ticket Hypothesis for Graph Neural Networks
- B^2SFL: A Bi-level Blockchained Architecture for Secure Federated Learning-based Traffic Prediction
- Vehicle Attribute Recognition by Appearance: Computer Vision Methods for Vehicle Type, Make and Model Classification
- Adaptive Neural Network-Based Approximation to Accelerate Eulerian Fluid Simulation
- Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient
- Continuous-in-Depth Neural Networks
- Shallow-UWnet : Compressed Model for Underwater Image Enhancement
- AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
- The model of an anomaly detector for HiLumi LHC magnets based on Recurrent Neural Networks and adaptive quantization
- AdaSpring: Context-adaptive and Runtime-evolutionary Deep Model Compression for Mobile Applications
- Optimally Scheduling CNN Convolutions for Efficient Memory Access
- Differentiable Neural Input Search for Recommender Systems
- Computer Vision Model Compression Techniques for Embedded Systems: A Survey
- On the Adversarial Robustness of Quantized Neural Networks
- Tensorial Neural Networks: Generalization of Neural Networks and Application to Model Compression
- Joint Iris Segmentation and Localization Using Deep Multi-task Learning Framework
- BraggNN: Fast X-ray Bragg Peak Analysis Using Deep Learning
- KTAN: Knowledge Transfer Adversarial Network
- The Search for Sparse, Robust Neural Networks
- CondenseNeXt: An Ultra-Efficient Deep Neural Network for Embedded Systems
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- OpenEI: An Open Framework for Edge Intelligence
- Restricted Recurrent Neural Networks
- Bayesian Deep Learning via Subnetwork Inference
- Constrained Deep Learning using Conditional Gradient and Applications in Computer Vision
- GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference
- Machine Learning for Cataract Classification and Grading on Ophthalmic Imaging Modalities: A Survey
- The Feasibility and Inevitability of Stealth Attacks
- Convolutional neural network compression for natural language processing
- Deep Face Recognition Model Compression via Knowledge Transfer and Distillation
- Deep Serial Number: Computational Watermarking for DNN Intellectual Property Protection
- Distilling BERT into Simple Neural Networks with Unlabeled Transfer Data
- Layer-wise Learning of Stochastic Neural Networks with Information Bottleneck
- Depth-wise Decomposition for Accelerating Separable Convolutions in Efficient Convolutional Neural Networks
- Wireless for Machine Learning
- Performance versus Complexity Study of Neural Network Equalizers in Coherent Optical Systems
- Neural 3D Scene Compression via Model Compression
- SpecNet: Spectral Domain Convolutional Neural Network
- PTEENet: Post-Trained Early-Exit Neural Networks Augmentation for Inference Cost Optimization
- DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression
- Transformed Regularization for Learning Sparse Deep Neural Networks
- Squeezed Deep 6DoF Object Detection Using Knowledge Distillation
- Learning from Higher-Layer Feature Visualizations
- PANDA: Facilitating Usable AI Development
- KD-MRI: A knowledge distillation framework for image reconstruction and image restoration in MRI workflow
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- PURSUhInT: In Search of Informative Hint Points Based on Layer Clustering for Knowledge Distillation
- Learning to Prune Deep Neural Networks via Reinforcement Learning
- HadaNets: Flexible Quantization Strategies for Neural Networks
- Deep Spiking Convolutional Neural Network for Single Object Localization Based On Deep Continuous Local Learning
- MWQ: Multiscale Wavelet Quantized Neural Networks
- To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference
- Deep Learning Towards Mobile Applications
- ESAI: Efficient Split Artificial Intelligence via Early Exiting Using Neural Architecture Search
- Accelerating Distributed ML Training via Selective Synchronization
- CDFI: Compression-Driven Network Design for Frame Interpolation
- Towards Self-Regulating AI: Challenges and Opportunities of AI Model Governance in Financial Services
- BiQGEMM: Matrix Multiplication with Lookup Table For Binary-Coding-based Quantized DNNs
- VEGA: Towards an End-to-End Configurable AutoML Pipeline
- Learning from a Teacher using Unlabeled Data
- 2-bit Model Compression of Deep Convolutional Neural Network on ASIC Engine for Image Retrieval
- Towards Stable Symbol Grounding with Zero-Suppressed State AutoEncoder
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- Towards a General Model of Knowledge for Facial Analysis by Multi-Source Transfer Learning
- Multi-head Knowledge Distillation for Model Compression
- P-KDGAN: Progressive Knowledge Distillation with GANs for One-class Novelty Detection
- Data Efficient Stagewise Knowledge Distillation
- 3DQ: Compact Quantized Neural Networks for Volumetric Whole Brain Segmentation
- B-DCGAN:Evaluation of Binarized DCGAN for FPGA
- EBJR: Energy-Based Joint Reasoning for Adaptive Inference
- Detecting Dead Weights and Units in Neural Networks
- Efficient Video Classification Using Fewer Frames
- Orthant Based Proximal Stochastic Gradient Method for -Regularized Optimization
- Neural Pruning via Growing Regularization
- ACDC: Weight Sharing in Atom-Coefficient Decomposed Convolution
- Data-Driven Compression of Convolutional Neural Networks
- ECG-DelNet: Delineation of Ambulatory Electrocardiograms with Mixed Quality Labeling Using Neural Networks
- Model compression for faster structural separation of macromolecules captured by Cellular Electron Cryo-Tomography
- Developing a Compressed Object Detection Model based on YOLOv4 for Deployment on Embedded GPU Platform of Autonomous System
- Benchmarking Inference Performance of Deep Learning Models on Analog Devices
- Compact representations of convolutional neural networks via weight pruning and quantization
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Cross-modal knowledge distillation for action recognition
- Deep Learning Techniques for Compressive Sensing-Based Reconstruction and Inference -- A Ubiquitous Systems Perspective
- Model based Multi-agent Reinforcement Learning with Tensor Decompositions
- Lossless Compression of Structured Convolutional Models via Lifting
- Parameter Prediction for Unseen Deep Architectures
- Line-Circle-Square (LCS): A Multilayered Geometric Filter for Edge-Based Detection
- AntMan: Sparse Low-Rank Compression to Accelerate RNN inference
- Stochastic Model Pruning via Weight Dropping Away and Back
- Adapt-and-Distill: Developing Small, Fast and Effective Pretrained Language Models for Domains
- iToF2dToF: A Robust and Flexible Representation for Data-Driven Time-of-Flight Imaging
- Going Beyond Classification Accuracy Metrics in Model Compression
- Beyond Preserved Accuracy: Evaluating Loyalty and Robustness of BERT Compression
- Deeplite Neutrino: An End-to-End Framework for Constrained Deep Learning Model Optimization
- Derivation and Analysis of Fast Bilinear Algorithms for Convolution
- ESPN: Extremely Sparse Pruned Networks
- A Unified Framework for Shot Type Classification Based on Subject Centric Lens
- A Survey on Large-scale Machine Learning
- CompNet: Neural networks growing via the compact network morphism
- Small, Accurate, and Fast Vehicle Re-ID on the Edge: the SAFR Approach
- BasisConv: A method for compressed representation and learning in CNNs
- BitHEP -- The Limits of Low-Precision ML in HEP
- Neural Networks at a Fraction with Pruned Quaternions
- Efficient Proximal Mapping of the 1-path-norm of Shallow Networks
- DiffPrune: Neural Network Pruning with Deterministic Approximate Binary Gates and Regularization
- Memory Requirement Reduction of Deep Neural Networks Using Low-bit Quantization of Parameters
- Efficient architecture for deep neural networks with heterogeneous sensitivity
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
- Deep Asymmetric Networks with a Set of Node-wise Variant Activation Functions
- A Survey on GAN Acceleration Using Memory Compression Technique
- The Elastic Lottery Ticket Hypothesis
- Localized Compression: Applying Convolutional Neural Networks to Compressed Images
- Residual-Guided In-Loop Filter Using Convolution Neural Network
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- I/O Lower Bounds for Auto-tuning of Convolutions in CNNs
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- Advancing biological super-resolution microscopy through deep learning: a brief review
- Finding Everything within Random Binary Networks
- AutoPruning for Deep Neural Network with Dynamic Channel Masking
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- Scaling Up Deep Neural Network Optimization for Edge Inference
- JIZHI: A Fast and Cost-Effective Model-As-A-Service System for Web-Scale Online Inference at Baidu
- Dense Pruning of Pointwise Convolutions in the Frequency Domain
- Reinforcement Learning in Factored Action Spaces using Tensor Decompositions
- Blind Adversarial Pruning: Balance Accuracy, Efficiency and Robustness
- Deep Learning Approximation: Zero-Shot Neural Network Speedup
- New Perspective on Progressive GANs Distillation for One-class Novelty Detection
- Gabor filter incorporated CNN for compression
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- Learning Realistic Patterns from Unrealistic Stimuli: Generalization and Data Anonymization
- Block-term Tensor Neural Networks
- Positioning yourself in the maze of Neural Text Generation: A Task-Agnostic Survey
- Novel Adaptive Binary Search Strategy-First Hybrid Pyramid- and Clustering-Based CNN Filter Pruning Method without Parameters Setting
- Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference
- A Step Towards Efficient Evaluation of Complex Perception Tasks in Simulation
- AUTOKD: Automatic Knowledge Distillation Into A Student Architecture Family
- Auto-Split: A General Framework of Collaborative Edge-Cloud AI
- Progressive Compressed Records: Taking a Byte out of Deep Learning Data
- Exact Backpropagation in Binary Weighted Networks with Group Weight Transformations
- Fast Intent Classification for Spoken Language Understanding
- Compressed Object Detection
- Iterative Training: Finding Binary Weight Deep Neural Networks with Layer Binarization
- PERMDNN: Efficient Compressed DNN Architecture with Permuted Diagonal Matrices
- A Domain-Oblivious Approach for Learning Concise Representations of Filtered Topological Spaces for Clustering
- Magnitude and Uncertainty Pruning Criterion for Neural Networks
- Neural networks adapting to datasets: learning network size and topology
- An Improving Framework of regularization for Network Compression
- A Highly Effective Low-Rank Compression of Deep Neural Networks with Modified Beam-Search and Modified Stable Rank
- Transferring Inter-Class Correlation
- Towards Modality Transferable Visual Information Representation with Optimal Model Compression
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- Development of Quantized DNN Library for Exact Hardware Emulation
- A Probabilistic Approach to Neural Network Pruning
- Nonlinear Tensor Ring Network
- Directed-Weighting Group Lasso for Eltwise Blocked CNN Pruning
- Simultaneously Learning Architectures and Features of Deep Neural Networks
- Decreasing the size of the Restricted Boltzmann machine