To prune, or not to prune: exploring the efficacy of pruning for model compression
arXiv:1710.01878
Abstract
Model pruning seeks to induce sparsity in a deep neural network's various connection matrices, thereby reducing the number of nonzero-valued parameters in the model. Recent reports (Han et al., 2015; Narang et al., 2017) prune deep networks at the cost of only a marginal loss in accuracy and achieve a sizable reduction in model size. This hints at the possibility that the baseline models in these experiments are perhaps severely over-parameterized at the outset and a viable alternative for model compression might be to simply reduce the number of hidden units while maintaining the model's dense connection structure, exposing a similar trade-off in model size and accuracy. We investigate these two distinct paths for model compression within the context of energy-efficient inference in resource-constrained environments and propose a new gradual pruning technique that is simple and straightforward to apply across a variety of models/datasets with minimal tuning and can be seamlessly incorporated within the training process. We compare the accuracy of large, but pruned models (large-sparse) and their smaller, but dense (small-dense) counterparts with identical memory footprint. Across a broad range of neural network architectures (deep CNNs, stacked LSTM, and seq2seq LSTM models), we find large-sparse models to consistently outperform small-dense models and achieve up to 10x reduction in number of non-zero parameters with minimal loss in accuracy.
Cited by in corpus (168)
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Machine Learning for Microcontroller-Class Hardware: A Review
- Integration of Neural Network-Based Symbolic Regression in Deep Learning for Scientific Discovery
- Sparse Networks from Scratch: Faster Training without Losing Performance
- QuantumNAS: Noise-Adaptive Search for Robust Quantum Circuits
- Adaptive Extreme Edge Computing for Wearable Devices
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
- Recurrent Neural Networks: An Embedded Computing Perspective
- Rigging the Lottery: Making All Tickets Winners
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
- Verification of Neural Network Behaviour: Formal Guarantees for Power System Applications
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- What Do Compressed Deep Neural Networks Forget?
- Performance and Complexity Analysis of bi-directional Recurrent Neural Network Models vs. Volterra Nonlinear Equalizers in Digital Coherent Systems
- Dynamic Model Pruning with Feedback
- Layer-adaptive sparsity for the Magnitude-based Pruning
- CAnet: Uplink-aided Downlink Channel Acquisition in FDD Massive MIMO using Deep Learning
- DAS-N2N: Machine learning Distributed Acoustic Sensing (DAS) signal denoising without clean data
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Soft-Demapping for Short Reach Optical Communication: A Comparison of Deep Neural Networks and Volterra Series
- Winning the Lottery with Continuous Sparsification
- Pruning Neural Networks at Initialization: Why are We Missing the Mark?
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Acceleration
- Hybrid Tensor Decomposition in Neural Network Compression
- DAIS: Automatic Channel Pruning via Differentiable Annealing Indicator Search
- PoPS: Policy Pruning and Shrinking for Deep Reinforcement Learning
- TinyissimoYOLO: A Quantized, Low-Memory Footprint, TinyML Object Detection Network for Low Power Microcontrollers
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- Movement Pruning: Adaptive Sparsity by Fine-Tuning
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- The Difficulty of Training Sparse Neural Networks
- Compressing RNNs for IoT devices by 15-38x using Kronecker Products
- Training independent subnetworks for robust prediction
- Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient
- Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient Implementation
- How fine can fine-tuning be? Learning efficient language models
- Compressing 3DCNNs Based on Tensor Train Decomposition
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- End-to-End Supermask Pruning: Learning to Prune Image Captioning Models
- Learned Threshold Pruning
- A Method for Medical Data Analysis Using the LogNNet for Clinical Decision Support Systems and Edge Computing in Healthcare
- Evaluating Single Event Upsets in Deep Neural Networks for Semantic Segmentation: an embedded system perspective
- Deep Neural Network Fingerprinting by Conferrable Adversarial Examples
- SeReNe: Sensitivity based Regularization of Neurons for Structured Sparsity in Neural Networks
- Anchor Pruning for Object Detection
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Sparse deep neural networks for modeling aluminum electrolysis dynamics
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Lookahead: A Far-Sighted Alternative of Magnitude-based Pruning
- Few Shot Network Compression via Cross Distillation
- Network Pruning That Matters: A Case Study on Retraining Variants
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- Selfish Sparse RNN Training
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- A hybrid inference system for improved curvature estimation in the level-set method using machine learning
- Accelerating Sparse Deep Neural Networks
- DASS: Differentiable Architecture Search for Sparse neural networks
- Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
- Muon trigger with fast Neural Networks on FPGA, a demonstrator
- Sanity-Checking Pruning Methods: Random Tickets can Win the Jackpot
- Finding Fast Transformers: One-Shot Neural Architecture Search by Component Composition
- FedDCT: Federated Learning of Large Convolutional Neural Networks on Resource Constrained Devices using Divide and Collaborative Training
- Pruning artificial neural networks: a way to find well-generalizing, high-entropy sharp minima
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Linear Mode Connectivity in Multitask and Continual Learning
- Ternary Hybrid Neural-Tree Networks for Highly Constrained IoT Applications
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural Networks
- Enabling Large Neural Networks on Tiny Microcontrollers with Swapping
- Compressing Deep Image Super-resolution Models
- Fast Vocabulary Transfer for Language Model Compression
- Sparse Weight Activation Training
- Adapting by Pruning: A Case Study on BERT
- SymbolNet: Neural Symbolic Regression with Adaptive Dynamic Pruning for Compression
- Sanity Checks for Lottery Tickets: Does Your Winning Ticket Really Win the Jackpot?
- TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning
- Improving Neural Network with Uniform Sparse Connectivity
- Sparse Systolic Tensor Array for Efficient CNN Hardware Acceleration
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Privacy-preserving Learning via Deep Net Pruning
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
- Wireless for Machine Learning
- Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm
- Fast and reliable uncertainty quantification with neural network ensembles for industrial image classification
- Chunked Autoregressive GAN for Conditional Waveform Synthesis
- Ordering Chaos: Memory-Aware Scheduling of Irregularly Wired Neural Networks for Edge Devices
- Communication-Computation Trade-Off in Resource-Constrained Edge Inference
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- A Spike in Performance: Training Hybrid-Spiking Neural Networks with Quantized Activation Functions
- Hessian-Aware Pruning and Optimal Neural Implant
- Clusterability in Neural Networks
- Robustness in Compressed Neural Networks for Object Detection
- Pruned Neural Networks are Surprisingly Modular
- Automated Model Compression by Jointly Applied Pruning and Quantization
- Network Pruning for Low-Rank Binary Indexing
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- Learning Language Specific Sub-network for Multilingual Machine Translation
- Synthesis and Pruning as a Dynamic Compression Strategy for Efficient Deep Neural Networks
- Compressing Language Models using Doped Kronecker Products
- Block-wise Dynamic Sparseness
- Dynamic Collective Intelligence Learning: Finding Efficient Sparse Model via Refined Gradients for Pruned Weights
- Image Captioning with Sparse Recurrent Neural Network
- PFGDF: Pruning Filter via Gaussian Distribution Feature for Deep Neural Networks Acceleration
- Pruning-then-Expanding Model for Domain Adaptation of Neural Machine Translation
- LCS: Learning Compressible Subspaces for Adaptive Network Compression at Inference Time
- Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)
- Dual-side Sparse Tensor Core
- Neural Networks at a Fraction with Pruned Quaternions
- Multi-word Tokenization for Sequence Compression
- An energy-based comparative analysis of common approaches to text classification in the Legal domain
- BWCP: Probabilistic Learning-to-Prune Channels for ConvNets via Batch Whitening
- A Bregman Learning Framework for Sparse Neural Networks
- Dynamic Sparsity Neural Networks for Automatic Speech Recognition
- Efficient Crowd Counting via Structured Knowledge Transfer
- Neural Machine Translation: A Review and Survey
- Accelerator-Aware Training for Transducer-Based Speech Recognition
- MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference
- The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together
- Search Spaces for Neural Model Training
- Adjoined Networks: A Training Paradigm with Applications to Network Compression
- Cross-Channel Intragroup Sparsity Neural Network
- Measure Twice, Cut Once: Quantifying Bias and Fairness in Deep Neural Networks
- Noisy Training Improves E2E ASR for the Edge
- The curious case of developmental BERTology: On sparsity, transfer learning, generalization and the brain
- SIPA: A Simple Framework for Efficient Networks
- The Role of Regularization in Shaping Weight and Node Pruning Dependency and Dynamics
- DiffPrune: Neural Network Pruning with Deterministic Approximate Binary Gates and Regularization
- Prune2Edge: A Multi-Phase Pruning Pipelines to Deep Ensemble Learning in IIoT
- SMOF: Squeezing More Out of Filters Yields Hardware-Friendly CNN Pruning
- Structured Pattern Pruning Using Regularization
- Bayesian Sparsification Methods for Deep Complex-valued Networks
- BERMo: What can BERT learn from ELMo?
- Neighbourhood Distillation: On the benefits of non end-to-end distillation
- MicroNet for Efficient Language Modeling
- Embedding Differentiable Sparsity into Deep Neural Network
- A Unified Speaker Adaptation Approach for ASR
- Comparative Study of Parameter Selection for Enhanced Edge Inference for a Multi-Output Regression model for Head Pose Estimation
- Induced Feature Selection by Structured Pruning
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation
- A Comparative Study of Neural Network Compression
- Neural Architecture Search via Bregman Iterations
- Dep-: Improving -based Network Sparsification via Dependency Modeling
- Importance-based Neuron Allocation for Multilingual Neural Machine Translation
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Modulating Regularization Frequency for Efficient Compression-Aware Model Training
- Optimization of DNN-based HSI Segmentation FPGA-based SoC for ADS: A Practical Approach
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- Extending Sparse Tensor Accelerators to Support Multiple Compression Formats
- Masked Training of Neural Networks with Partial Gradients
- DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths
- Cascade Weight Shedding in Deep Neural Networks: Benefits and Pitfalls for Network Pruning
- GHFP: Gradually Hard Filter Pruning
- How Well Do Sparse Imagenet Models Transfer?
- Learning Pruned Structure and Weights Simultaneously from Scratch: an Attention based Approach
- How Compact?: Assessing Compactness of Representations through Layer-Wise Pruning
- Pruning Attention Heads of Transformer Models Using A* Search: A Novel Approach to Compress Big NLP Architectures
- Low-Rank+Sparse Tensor Compression for Neural Networks
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- Does a sparse ReLU network training problem always admit an optimum?