Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
arXiv:1607.03250
Abstract
State-of-the-art neural networks are getting deeper and wider. While their performance increases with the increasing number of layers and neurons, it is crucial to design an efficient deep architecture in order to reduce computational and memory costs. Designing an efficient neural network, however, is labor intensive requiring many experiments, and fine-tunings. In this paper, we introduce network trimming which iteratively optimizes the network by pruning unimportant neurons based on analysis of their outputs on a large dataset. Our algorithm is inspired by an observation that the outputs of a significant portion of neurons in a large network are mostly zero, regardless of what inputs the network received. These zero activation neurons are redundant, and can be removed without affecting the overall accuracy of the network. After pruning the zero activation neurons, we retrain the network using the weights before pruning as initialization. We alternate the pruning and retraining to further reduce zero activations in a network. Our experiments on the LeNet and VGG-16 show that we can achieve high compression ratio of parameters without losing or even achieving higher accuracy than the original network.
References in corpus (5)
Cited by in corpus (177)
- Pruning Convolutional Neural Networks for Resource Efficient Inference
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- Regularized Deep Networks in Intelligent Transportation Systems: A Taxonomy and a Case Study
- Channel Pruning for Accelerating Very Deep Neural Networks
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- Recent Advances in Convolutional Neural Networks
- Bringing AI To Edge: From Deep Learning's Perspective
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Discrimination-aware Network Pruning for Deep Model Compression
- An Entropy-based Pruning Method for CNN Compression
- Stabilizing the Lottery Ticket Hypothesis
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
- Pruning Deep Convolutional Neural Networks Architectures with Evolution Strategy
- MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning
- The Power of Sparsity in Convolutional Neural Networks
- Continual Learning via Neural Pruning
- Proving the Lottery Ticket Hypothesis: Pruning is All You Need
- NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
- A novel channel pruning method for deep neural network compression
- Filter Pruning by Switching to Neighboring CNNs with Good Attributes
- Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon
- Cluster Pruning: An Efficient Filter Pruning Method for Edge AI Vision Applications
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- PruneTrain: Fast Neural Network Training by Dynamic Sparse Model Reconfiguration
- Neural Network Distiller: A Python Package For DNN Compression Research
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- NeST: A Neural Network Synthesis Tool Based on a Grow-and-Prune Paradigm
- HRel: Filter Pruning based on High Relevance between Activation Maps and Class Labels
- 2PFPCE: Two-Phase Filter Pruning Based on Conditional Entropy
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- Pruning Algorithms to Accelerate Convolutional Neural Networks for Edge Applications: A Survey
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
- PoPS: Policy Pruning and Shrinking for Deep Reinforcement Learning
- You Only Search Once: Single Shot Neural Architecture Search via Direct Sparse Optimization
- DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic Framework
- Layer-compensated Pruning for Resource-constrained Convolutional Neural Networks
- Gradual Channel Pruning while Training using Feature Relevance Scores for Convolutional Neural Networks
- Accelerating Neural ODEs Using Model Order Reduction
- Multi-Task Zipping via Layer-wise Neuron Sharing
- Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
- Data-Driven Sparse Structure Selection for Deep Neural Networks
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Task dependent Deep LDA pruning of neural networks
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices
- Joint Multi-Dimension Pruning via Numerical Gradient Update
- Channel Pruning via Automatic Structure Search
- A Programmable Approach to Neural Network Compression
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Computer Vision Model Compression Techniques for Embedded Systems: A Survey
- Knapsack Pruning with Inner Distillation
- Data-Independent Neural Pruning via Coresets
- Modeling of Pruning Techniques for Deep Neural Networks Simplification
- Cooperative data-driven modeling
- DMCP: Differentiable Markov Channel Pruning for Neural Networks
- One Person, One Model, One World: Learning Continual User Representation without Forgetting
- AdderNet: Do We Really Need Multiplications in Deep Learning?
- Convolutional Neural Network Pruning with Structural Redundancy Reduction
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Neural network relief: a pruning algorithm based on neural activity
- Shift-based Primitives for Efficient Convolutional Neural Networks
- CoCoPIE: Making Mobile AI Sweet As PIE --Compression-Compilation Co-Design Goes a Long Way
- Towards Evolutional Compression
- Understanding Convolutional Neural Networks with Information Theory: An Initial Exploration
- BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
- HRank: Filter Pruning using High-Rank Feature Map
- Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- PruneNet: Channel Pruning via Global Importance
- Taxonomy of Saliency Metrics for Channel Pruning
- Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
- Feature Flow Regularization: Improving Structured Sparsity in Deep Neural Networks
- Efficient Inference of CNNs via Channel Pruning
- CUP: Cluster Pruning for Compressing Deep Neural Networks
- Stochastic Channel-Based Federated Learning for Medical Data Privacy Preserving
- Depth-wise Decomposition for Accelerating Separable Convolutions in Efficient Convolutional Neural Networks
- Structured Pruning for Efficient ConvNets via Incremental Regularization
- FireNet: Real-time Segmentation of Fire Perimeter from Aerial Video
- GAN Slimming: All-in-One GAN Compression by A Unified Optimization Framework
- Toward Compact Deep Neural Networks via Energy-Aware Pruning
- Exploiting Kernel Sparsity and Entropy for Interpretable CNN Compression
- Full-Cycle Energy Consumption Benchmark for Low-Carbon Computer Vision
- Composition of Saliency Metrics for Channel Pruning with a Myopic Oracle
- Pruning Deep Neural Networks using Partial Least Squares
- ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- Parsimonious Inference on Convolutional Neural Networks: Learning and applying on-line kernel activation rules
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Towards Compact CNNs via Collaborative Compression
- Pruning at a Glance: Global Neural Pruning for Model Compression
- Demystifying Neural Network Filter Pruning
- Training Compact CNNs for Image Classification using Dynamic-coded Filter Fusion
- Attribution Preservation in Network Compression for Reliable Network Interpretation
- Convolutional Neural Network Simplification with Progressive Retraining
- An Image Enhancing Pattern-based Sparsity for Real-time Inference on Mobile Devices
- Neural Network Compression Via Sparse Optimization
- Shapley Value as Principled Metric for Structured Network Pruning
- Avoiding overfitting of multilayer perceptrons by training derivatives
- Structural Pruning in Deep Neural Networks: A Small-World Approach
- Assessing Intelligence in Artificial Neural Networks
- Filter Sketch for Network Pruning
- RARTS: An Efficient First-Order Relaxed Architecture Search Method
- Privacy Preserving Stochastic Channel-Based Federated Learning with Neural Network Pruning
- Spectral Pruning: Compressing Deep Neural Networks via Spectral Analysis and its Generalization Error
- CALPA-NET: Channel-pruning-assisted Deep Residual Network for Steganalysis of Digital Images
- Pufferfish: Communication-efficient Models At No Extra Cost
- FlexSA: Flexible Systolic Array Architecture for Efficient Pruned DNN Model Training
- Network Pruning using Adaptive Exemplar Filters
- Do We Need Fully Connected Output Layers in Convolutional Networks?
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- Synthesis and Pruning as a Dynamic Compression Strategy for Efficient Deep Neural Networks
- Data-Independent Structured Pruning of Neural Networks via Coresets
- Parameter Efficient Deep Neural Networks with Bilinear Projections
- Baseline Pruning-Based Approach to Trojan Detection in Neural Networks
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- RePr: Improved Training of Convolutional Filters
- Principal Component Networks: Parameter Reduction Early in Training
- AdaptCL: Efficient Collaborative Learning with Dynamic and Adaptive Pruning
- Line-Circle-Square (LCS): A Multilayered Geometric Filter for Edge-Based Detection
- HALP: Hardware-Aware Latency Pruning
- SASL: Saliency-Adaptive Sparsity Learning for Neural Network Acceleration
- Joint Channel and Weight Pruning for Model Acceleration on Moblie Devices
- Universal Adder Neural Networks
- Supervised Robustness-preserving Data-free Neural Network Pruning
- Why Lottery Ticket Wins? A Theoretical Perspective of Sample Complexity on Pruned Neural Networks
- Going Beyond Classification Accuracy Metrics in Model Compression
- AdaPruner: Adaptive Channel Pruning and Effective Weights Inheritance
- Training Sparse Neural Networks using Compressed Sensing
- Filter Pruning using Hierarchical Group Sparse Regularization for Deep Convolutional Neural Networks
- Content-Aware GAN Compression
- Thanks for Nothing: Predicting Zero-Valued Activations with Lightweight Convolutional Neural Networks
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- Exploiting Weight Redundancy in CNNs: Beyond Pruning and Quantization
- UCP: Uniform Channel Pruning for Deep Convolutional Neural Networks Compression and Acceleration
- Localization-aware Channel Pruning for Object Detection
- Automatic Inference of Cross-modal Connection Topologies for X-CNNs
- Pruning Ternary Quantization
- Rethinking Convolutional Features in Correlation Filter Based Tracking
- A Layer Decomposition-Recomposition Framework for Neuron Pruning towards Accurate Lightweight Networks
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- Improving Network Slimming with Nonconvex Regularization
- Implicit Filter Sparsification In Convolutional Neural Networks
- Pruning with Compensation: Efficient Channel Pruning for Deep Convolutional Neural Networks
- Channel Pruning Guided by Classification Loss and Feature Importance
- Class-dependent Compression of Deep Neural Networks
- Automatic Neural Network Pruning that Efficiently Preserves the Model Accuracy
- Channel-wise pruning of neural networks with tapering resource constraint
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- A One-step Pruning-recovery Framework for Acceleration of Convolutional Neural Networks
- Carrying out CNN Channel Pruning in a White Box
- Alternate Model Growth and Pruning for Efficient Training of Recommendation Systems
- Self-grouping Convolutional Neural Networks
- Online Filter Clustering and Pruning for Efficient Convnets
- EZCrop: Energy-Zoned Channels for Robust Output Pruning
- Content-Aware Convolutional Neural Networks
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- A Probabilistic Approach to Neural Network Pruning
- Blending Pruning Criteria for Convolutional Neural Networks
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- Network Pruning via Annealing and Direct Sparsity Control
- Dep-: Improving -based Network Sparsification via Dependency Modeling
- Speeding up convolutional networks pruning with coarse ranking
- Cluster Regularized Quantization for Deep Networks Compression
- Cogradient Descent for Dependable Learning
- Out-of-the-box channel pruned networks
- Light Multi-segment Activation for Model Compression
- Feature Statistics Guided Efficient Filter Pruning
- Pruning-Aware Merging for Efficient Multitask Inference
- Scientific Calculator for Designing Trojan Detectors in Neural Networks
- Weight Evolution: Improving Deep Neural Networks Training through Evolving Inferior Weight Values
- Reliable Identification of Redundant Kernels for Convolutional Neural Network Compression
- Channel selection using Gumbel Softmax
- Sparsity-Control Ternary Weight Networks
- Efficient Structured Pruning and Architecture Searching for Group Convolution
- GRIM: A General, Real-Time Deep Learning Inference Framework for Mobile Devices based on Fine-Grained Structured Weight Sparsity