To prune, or not to prune: exploring the efficacy of pruning for model compression
arXiv:1710.01878
Abstract
Model pruning seeks to induce sparsity in a deep neural network's various connection matrices, thereby reducing the number of nonzero-valued parameters in the model. Recent reports (Han et al., 2015; Narang et al., 2017) prune deep networks at the cost of only a marginal loss in accuracy and achieve a sizable reduction in model size. This hints at the possibility that the baseline models in these experiments are perhaps severely over-parameterized at the outset and a viable alternative for model compression might be to simply reduce the number of hidden units while maintaining the model's dense connection structure, exposing a similar trade-off in model size and accuracy. We investigate these two distinct paths for model compression within the context of energy-efficient inference in resource-constrained environments and propose a new gradual pruning technique that is simple and straightforward to apply across a variety of models/datasets with minimal tuning and can be seamlessly incorporated within the training process. We compare the accuracy of large, but pruned models (large-sparse) and their smaller, but dense (small-dense) counterparts with identical memory footprint. Across a broad range of neural network architectures (deep CNNs, stacked LSTM, and seq2seq LSTM models), we find large-sparse models to consistently outperform small-dense models and achieve up to 10x reduction in number of non-zero parameters with minimal loss in accuracy.
Cited by in corpus (57)
- Dynamic Model Pruning with Feedback
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Hybrid Tensor Decomposition in Neural Network Compression
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- How fine can fine-tuning be? Learning efficient language models
- A Method for Medical Data Analysis Using the LogNNet for Clinical Decision Support Systems and Edge Computing in Healthcare
- End-to-End Supermask Pruning: Learning to Prune Image Captioning Models
- Lookahead: A Far-Sighted Alternative of Magnitude-based Pruning
- Accelerating Sparse Deep Neural Networks
- Finding Fast Transformers: One-Shot Neural Architecture Search by Component Composition
- Ternary Hybrid Neural-Tree Networks for Highly Constrained IoT Applications
- TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning
- Sanity Checks for Lottery Tickets: Does Your Winning Ticket Really Win the Jackpot?
- Improving Neural Network with Uniform Sparse Connectivity
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Privacy-preserving Learning via Deep Net Pruning
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Ordering Chaos: Memory-Aware Scheduling of Irregularly Wired Neural Networks for Edge Devices
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Clusterability in Neural Networks
- Network Pruning for Low-Rank Binary Indexing
- Automated Model Compression by Jointly Applied Pruning and Quantization
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
- Synthesis and Pruning as a Dynamic Compression Strategy for Efficient Deep Neural Networks
- Pruning-then-Expanding Model for Domain Adaptation of Neural Machine Translation
- Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)
- LCS: Learning Compressible Subspaces for Adaptive Network Compression at Inference Time
- Dual-side Sparse Tensor Core
- Measure Twice, Cut Once: Quantifying Bias and Fairness in Deep Neural Networks
- Accelerator-Aware Training for Transducer-Based Speech Recognition
- The Role of Regularization in Shaping Weight and Node Pruning Dependency and Dynamics
- Search Spaces for Neural Model Training
- The curious case of developmental BERTology: On sparsity, transfer learning, generalization and the brain
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together
- DiffPrune: Neural Network Pruning with Deterministic Approximate Binary Gates and Regularization
- Noisy Training Improves E2E ASR for the Edge
- A Comparative Study of Neural Network Compression
- Dep-: Improving -based Network Sparsification via Dependency Modeling
- Neural Architecture Search via Bregman Iterations
- How Compact?: Assessing Compactness of Representations through Layer-Wise Pruning
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- Extending Sparse Tensor Accelerators to Support Multiple Compression Formats
- Cascade Weight Shedding in Deep Neural Networks: Benefits and Pitfalls for Network Pruning
- GHFP: Gradually Hard Filter Pruning
- Low-Rank+Sparse Tensor Compression for Neural Networks
- MicroNet for Efficient Language Modeling
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- SMOF: Squeezing More Out of Filters Yields Hardware-Friendly CNN Pruning
- BERMo: What can BERT learn from ELMo?
- A Unified Speaker Adaptation Approach for ASR
- Embedding Differentiable Sparsity into Deep Neural Network
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation
- Structured Pattern Pruning Using Regularization
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Importance-based Neuron Allocation for Multilingual Neural Machine Translation