Pruning Convolutional Neural Networks for Resource Efficient Inference
arXiv:1611.06440
Abstract
We propose a new formulation for pruning convolutional kernels in neural networks to enable efficient inference. We interleave greedy criteria-based pruning with fine-tuning by backpropagation - a computationally efficient procedure that maintains good generalization in the pruned network. We propose a new criterion based on Taylor expansion that approximates the change in the cost function induced by pruning network parameters. We focus on transfer learning, where large pretrained networks are adapted to specialized tasks. The proposed criterion demonstrates superior performance compared to other criteria, e.g. the norm of kernel weights or feature map activation, for pruning large CNNs after adaptation to fine-grained classification tasks (Birds-200 and Flowers-102) relaying only on the first order gradient information. We also show that pruning can lead to more than 10x theoretical (5x practical) reduction in adapted 3D-convolutional filters with a small drop in accuracy in a recurrent gesture classifier. Finally, we show results for the large-scale ImageNet dataset to emphasize the flexibility of our approach.
17 pages, 14 figures, ICLR 2017 paper
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Learning with Limited Numerical Precision
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Learning Structured Sparsity in Deep Neural Networks
- Bird Species Categorization Using Pose Normalized Deep Convolutional Nets
- maxDNN: An Efficient Convolution Kernel for Deep Learning with Maxwell GPUs
Cited by in corpus (271)
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- A Review on Deep Learning Techniques Applied to Semantic Segmentation
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- Slimmable Neural Networks
- Self-Supervised Model Adaptation for Multimodal Semantic Segmentation
- Ablation Studies in Artificial Neural Networks
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- Graph-based Spatial-temporal Feature Learning for Neuromorphic Vision Sensing
- XONN: XNOR-based Oblivious Deep Neural Network Inference
- Discrimination-aware Network Pruning for Deep Model Compression
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Edge Intelligence: Architectures, Challenges, and Applications
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- XNOR-Net++: Improved Binary Neural Networks
- Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
- Joint Device-Edge Inference over Wireless Links with Pruning
- Pruning Deep Convolutional Neural Networks Architectures with Evolution Strategy
- Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP
- Accelerator-Aware Pruning for Convolutional Neural Networks
- Dual Dynamic Inference: Enabling More Efficient, Adaptive and Controllable Deep Inference
- A Systematic Review on Model Watermarking for Neural Networks
- Deep Learning-based Implicit CSI Feedback in Massive MIMO
- Environmental Sound Classification on the Edge: A Pipeline for Deep Acoustic Networks on Extremely Resource-Constrained Devices
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- What is the State of Neural Network Pruning?
- A Low Effort Approach to Structured CNN Design Using PCA
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Learning Sparse Networks Using Targeted Dropout
- NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
- RepMLP: Re-parameterizing Convolutions into Fully-connected Layers for Image Recognition
- EDVR: Video Restoration with Enhanced Deformable Convolutional Networks
- Cluster Pruning: An Efficient Filter Pruning Method for Edge AI Vision Applications
- AutoTune: Automatically Tuning Convolutional Neural Networks for Improved Transfer Learning
- Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation
- Crossbar-aware neural network pruning
- Is Attention Better Than Matrix Decomposition?
- Neural Network Distiller: A Python Package For DNN Compression Research
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Video Semantic Segmentation with Distortion-Aware Feature Correction
- Attention-Based 3D Seismic Fault Segmentation Training by a Few 2D Slice Labels
- An Experimental Study of Reduced-Voltage Operation in Modern FPGAs for Neural Network Acceleration
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Batch-Shaping for Learning Conditional Channel Gated Networks
- 2PFPCE: Two-Phase Filter Pruning Based on Conditional Entropy
- Are Sixteen Heads Really Better than One?
- CARAFE: Content-Aware ReAssembly of FEatures
- LFFD: A Light and Fast Face Detector for Edge Devices
- M2KD: Multi-model and Multi-level Knowledge Distillation for Incremental Learning
- Pruning Algorithms to Accelerate Convolutional Neural Networks for Edge Applications: A Survey
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- Progressive Skeletonization: Trimming more fat from a network at initialization
- TinyissimoYOLO: A Quantized, Low-Memory Footprint, TinyML Object Detection Network for Low Power Microcontrollers
- Fingerprint Presentation Attack Detector Using Global-Local Model
- Knowledge Distillation via Route Constrained Optimization
- When BERT Plays the Lottery, All Tickets Are Winning
- On-device Training: A First Overview on Existing Systems
- HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision
- AutoPruner: An End-to-End Trainable Filter Pruning Method for Efficient Deep Model Inference
- Pruning vs XNOR-Net: A Comprehensive Study of Deep Learning for Audio Classification on Edge-devices
- AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted Surgery
- Accelerating Neural ODEs Using Model Order Reduction
- Neural Networks Compression for Language Modeling
- SCSP: Spectral Clustering Filter Pruning with Soft Self-adaption Manners
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- Training independent subnetworks for robust prediction
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Understanding Neural Networks and Individual Neuron Importance via Information-Ordered Cumulative Ablation
- PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training
- Non-Structured DNN Weight Pruning -- Is It Beneficial in Any Platform?
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Channel Compression: Rethinking Information Redundancy among Channels in CNN Architecture
- Adaptive Neural Network-Based Approximation to Accelerate Eulerian Fluid Simulation
- Preferences Prediction using a Gallery of Mobile Device based on Scene Recognition and Object Detection
- Uncertainty-guided Continual Learning with Bayesian Neural Networks
- Greedy Layerwise Learning Can Scale to ImageNet
- Inference skipping for more efficient real-time speech enhancement with parallel RNNs
- Role of Data Augmentation Strategies in Knowledge Distillation for Wearable Sensor Data
- A Method for Medical Data Analysis Using the LogNNet for Clinical Decision Support Systems and Edge Computing in Healthcare
- A Closer Look at Structured Pruning for Neural Network Compression
- Distilled Neural Networks for Efficient Learning to Rank
- A Programmable Approach to Neural Network Compression
- Accelerate CNNs from Three Dimensions: A Comprehensive Pruning Framework
- Dissecting FLOPs along input dimensions for GreenAI cost estimations
- Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- Structural Compression of Convolutional Neural Networks
- STEERAGE: Synthesis of Neural Networks Using Architecture Search and Grow-and-Prune Methods
- A Gradient Flow Framework For Analyzing Network Pruning
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Knapsack Pruning with Inner Distillation
- Towards Timely Video Analytics Services at the Network Edge
- HAWQV3: Dyadic Neural Network Quantization
- Network Pruning That Matters: A Case Study on Retraining Variants
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Selfish Sparse RNN Training
- Can Subnetwork Structure be the Key to Out-of-Distribution Generalization?
- I-BERT: Integer-only BERT Quantization
- Supervised machine learning to estimate instabilities in chaotic systems: estimation of local Lyapunov exponents
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- LeanConvNets: Low-cost Yet Effective Convolutional Neural Networks
- DMCP: Differentiable Markov Channel Pruning for Neural Networks
- SiPPing Neural Networks: Sensitivity-informed Provable Pruning of Neural Networks
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- Multilingual Byte2Speech Models for Scalable Low-resource Speech Synthesis
- Copernicus: Characterizing the Performance Implications of Compression Formats Used in Sparse Workloads
- VESR-Net: The Winning Solution to Youku Video Enhancement and Super-Resolution Challenge
- Convolutional Neural Network Pruning with Structural Redundancy Reduction
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Accelerating Sparse Deep Neural Networks
- Weakly-supervised Learning For Catheter Segmentation in 3D Frustum Ultrasound
- Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
- Pruning via Iterative Ranking of Sensitivity Statistics
- Benchmarking Quantized Neural Networks on FPGAs with FINN
- BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
- CoCoPIE: Making Mobile AI Sweet As PIE --Compression-Compilation Co-Design Goes a Long Way
- Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity
- Under the Hood of Neural Networks: Characterizing Learned Representations by Functional Neuron Populations and Network Ablations
- MECKD: Deep Learning-Based Fall Detection in Multilayer Mobile Edge Computing With Knowledge Distillation
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Directional Pruning of Deep Neural Networks
- Logic Shrinkage: Learned FPGA Netlist Sparsity for Efficient Neural Network Inference
- GeneCAI: Genetic Evolution for Acquiring Compact AI
- Hardware-Efficient Photonic Tensor Core: Accelerating Deep Neural Networks with Structured Compression
- HRank: Filter Pruning using High-Rank Feature Map
- Equivalent and Approximate Transformations of Deep Neural Networks
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- Convolutional Neural Network based Multiple-Rate Compressive Sensing for Massive MIMO CSI Feedback: Design, Simulation, and Analysis
- Deep Hashing with Category Mask for Fast Video Retrieval
- Taxonomy of Saliency Metrics for Channel Pruning
- Sparse Weight Activation Training
- Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
- VEDLIoT -- Next generation accelerated AIoT systems and applications
- Efficient Inference of CNNs via Channel Pruning
- SQuantizer: Simultaneous Learning for Both Sparse and Low-precision Neural Networks
- A Novel Correlation-optimized Deep Learning Method for Wind Speed Forecast
- Exploiting Kernel Sparsity and Entropy for Interpretable CNN Compression
- Toward Compact Deep Neural Networks via Energy-Aware Pruning
- Bayesian Sparse learning with preconditioned stochastic gradient MCMC and its applications
- Topological Persistence Guided Knowledge Distillation for Wearable Sensor Data
- Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization
- Convolutional Fully-Connected Capsule Network (CFC-CapsNet): A Novel and Fast Capsule Network
- Composition of Saliency Metrics for Channel Pruning with a Myopic Oracle
- Fast Neural Architecture Construction using EnvelopeNets
- Towards Compact CNNs via Collaborative Compression
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
- You are caught stealing my winning lottery ticket! Making a lottery ticket claim its ownership
- Alleviating the Inequality of Attention Heads for Neural Machine Translation
- An Image Enhancing Pattern-based Sparsity for Real-time Inference on Mobile Devices
- Backdoor Attacks on Federated Learning with Lottery Ticket Hypothesis
- Learning Strict Identity Mappings in Deep Residual Networks
- A Markovian Model-Driven Deep Learning Framework for Massive MIMO CSI Feedback
- GreenFlow: A Computation Allocation Framework for Building Environmentally Sound Recommendation System
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Semantic Communication Systems for Speech Transmission
- COP: Customized Deep Model Compression via Regularized Correlation-Based Filter-Level Pruning
- Training Compact CNNs for Image Classification using Dynamic-coded Filter Fusion
- Convolutional Neural Network Simplification with Progressive Retraining
- A New Compensatory Genetic Algorithm-Based Method for Effective Compressed Multi-function Convolutional Neural Network Model Selection with Multi-Objective Optimization
- Mixed-Precision Quantized Neural Network with Progressively Decreasing Bitwidth For Image Classification and Object Detection
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- Hessian-Aware Pruning and Optimal Neural Implant
- Sparse Flows: Pruning Continuous-depth Models
- Single-shot Channel Pruning Based on Alternating Direction Method of Multipliers
- AI-oriented Medical Workload Allocation for Hierarchical Cloud/Edge/Device Computing
- Online Knowledge Distillation via Multi-branch Diversity Enhancement
- Neural Network Based Optimization of Transmit Beamforming and RIS Coefficients Using Channel Covariances in MISO Downlink
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise Decomposition
- Lottery Jackpots Exist in Pre-trained Models
- Scaling Up Exact Neural Network Compression by ReLU Stability
- WAFFLE: Watermarking in Federated Learning
- Overcoming Catastrophic Forgetting by Neuron-level Plasticity Control
- Dynamic Sparse Graph for Efficient Deep Learning
- Distilling with Performance Enhanced Students
- Enhancing the Regularization Effect of Weight Pruning in Artificial Neural Networks
- Correlation Congruence for Knowledge Distillation
- Deep Learning-Based CSI Feedback for Beamforming in Single- and Multi-cell Massive MIMO Systems
- Deleting object selective units in a fully-connected layer of deep convolutional networks improves classification performance
- Exploiting the Full Capacity of Deep Neural Networks while Avoiding Overfitting by Targeted Sparsity Regularization
- Pruning Convolutional Neural Networks for Image Instance Retrieval
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- Baseline Pruning-Based Approach to Trojan Detection in Neural Networks
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- Network Automatic Pruning: Start NAP and Take a Nap
- Identifying Critical Neurons in ANN Architectures using Mixed Integer Programming
- When to Prune? A Policy towards Early Structural Pruning
- Single-Net Continual Learning with Progressive Segmented Training (PST)
- HALP: Hardware-Aware Latency Pruning
- FALCON: Lightweight and Accurate Convolution
- AdaptCL: Efficient Collaborative Learning with Dynamic and Adaptive Pruning
- A novel method for identifying the deep neural network model with the Serial Number
- LeanResNet: A Low-cost Yet Effective Convolutional Residual Networks
- Distributed Learning on Heterogeneous Resource-Constrained Devices
- RicciNets: Curvature-guided Pruning of High-performance Neural Networks Using Ricci Flow
- Plug-in, Trainable Gate for Streamlining Arbitrary Neural Networks
- DECORE: Deep Compression with Reinforcement Learning
- Balancing Specialization, Generalization, and Compression for Detection and Tracking
- Prior Activation Distribution (PAD): A Versatile Representation to Utilize DNN Hidden Units
- Channel Pruning via Optimal Thresholding
- Going Beyond Classification Accuracy Metrics in Model Compression
- DASNet: Dynamic Activation Sparsity for Neural Network Efficiency Improvement
- Stochastic In-Face Frank-Wolfe Methods for Non-Convex Optimization and Sparse Neural Network Training
- Supervised Robustness-preserving Data-free Neural Network Pruning
- Filter Pruning using Hierarchical Group Sparse Regularization for Deep Convolutional Neural Networks
- Leveraging Structured Pruning of Convolutional Neural Networks
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- TOCO: A Framework for Compressing Neural Network Models Based on Tolerance Analysis
- Batch Normalization Sampling
- SIPA: A Simple Framework for Efficient Networks
- Channel-wise pruning of neural networks with tapering resource constraint
- Self-Reorganizing and Rejuvenating CNNs for Increasing Model Capacity Utilization
- Group and Exclusive Sparse Regularization-based Continual Learning of CNNs
- CompactNet: Platform-Aware Automatic Optimization for Convolutional Neural Networks
- Grassmannian Packings in Neural Networks: Learning with Maximal Subspace Packings for Diversity and Anti-Sparsity
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- ASCAI: Adaptive Sampling for acquiring Compact AI
- Stitching for Neuroevolution: Recombining Deep Neural Networks without Breaking Them
- Model Compression for Resource-Constrained Mobile Robots
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Spatial Sharing of GPU for Autotuning DNN models
- Digging Deeper into CRNN Model in Chinese Text Images Recognition
- Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization
- NeuralScale: Efficient Scaling of Neurons for Resource-Constrained Deep Neural Networks
- Correcting Momentum in Temporal Difference Learning
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- Dynamic Multi-path Neural Network
- Domain Adaptation Regularization for Spectral Pruning
- Rapid Elastic Architecture Search under Specialized Classes and Resource Constraints
- Channel Pruning Guided by Classification Loss and Feature Importance
- Implicit Filter Sparsification In Convolutional Neural Networks
- Neural networks adapting to datasets: learning network size and topology
- Carrying out CNN Channel Pruning in a White Box
- Alternate Model Growth and Pruning for Efficient Training of Recommendation Systems
- Modulating Regularization Frequency for Efficient Compression-Aware Model Training
- EZCrop: Energy-Zoned Channels for Robust Output Pruning
- Transfer Learning with Binary Neural Networks
- MoRS: An Approximate Fault Modelling Framework for Reduced-Voltage SRAMs
- How Well Do Sparse Imagenet Models Transfer?
- Plant 'n' Seek: Can You Find the Winning Ticket?
- A Survey on Green Deep Learning
- RGP: Neural Network Pruning through Its Regular Graph Structure
- Neural Weight Step Video Compression
- Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation
- Induced Feature Selection by Structured Pruning
- Putting 3D Spatially Sparse Networks on a Diet
- Cogradient Descent for Dependable Learning
- Blending Pruning Criteria for Convolutional Neural Networks
- Student Helping Teacher: Teacher Evolution via Self-Knowledge Distillation
- An Adaptive Empirical Bayesian Method for Sparse Deep Learning
- Channel Planting for Deep Neural Networks using Knowledge Distillation
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- Joint Regularization on Activations and Weights for Efficient Neural Network Pruning
- HALO: Learning to Prune Neural Networks with Shrinkage
- AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
- FastSal: a Computationally Efficient Network for Visual Saliency Prediction
- Channel selection using Gumbel Softmax
- Self-grouping Convolutional Neural Networks
- Architecture Compression
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- DAC: Data-free Automatic Acceleration of Convolutional Networks
- How Compact?: Assessing Compactness of Representations through Layer-Wise Pruning
- Speeding up convolutional networks pruning with coarse ranking
- Investigating Channel Pruning through Structural Redundancy Reduction -- A Statistical Study
- Online Filter Clustering and Pruning for Efficient Convnets
- Block-term Tensor Neural Networks
- Activation Map Adaptation for Effective Knowledge Distillation
- Grow-Push-Prune: aligning deep discriminants for effective structural network compression
- LEAN: graph-based pruning for convolutional neural networks by extracting longest chains