Learning both Weights and Connections for Efficient Neural Networks
arXiv:1506.02626
Abstract
Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems. Also, conventional networks fix the architecture before training starts; as a result, training cannot improve the architecture. To address these limitations, we describe a method to reduce the storage and computation required by neural networks by an order of magnitude without affecting their accuracy by learning only the important connections. Our method prunes redundant connections using a three-step method. First, we train the network to learn which connections are important. Next, we prune the unimportant connections. Finally, we retrain the network to fine tune the weights of the remaining connections. On the ImageNet dataset, our method reduced the number of parameters of AlexNet by a factor of 9x, from 61 million to 6.7 million, without incurring accuracy loss. Similar experiments with VGG-16 found that the number of parameters can be reduced by 13x, from 138 million to 10.3 million, again with no loss of accuracy.
Published as a conference paper at NIPS 2015
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Memory Bounded Deep Convolutional Networks
- Deep Fried Convnets
Cited by in corpus (56)
- Knowledge Distillation: A Survey
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- Embedding Watermarks into Deep Neural Networks
- Towards Explainable Artificial Intelligence
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Large-Scale Evolution of Image Classifiers
- A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers
- Variational Dropout Sparsifies Deep Neural Networks
- Supporting Very Large Models using Automatic Dataflow Graph Partitioning
- FETCH: A deep-learning based classifier for fast transient classification
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- MaskConnect: Connectivity Learning by Gradient Descent
- Constructing Energy-efficient Mixed-precision Neural Networks through Principal Component Analysis for Edge Intelligence
- Chess AI: Competing Paradigms for Machine Intelligence
- FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
- MARS: Multi-macro Architecture SRAM CIM-Based Accelerator with Co-designed Compressed Neural Networks
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- Fast On-the-fly Retraining-free Sparsification of Convolutional Neural Networks
- -ARM: Network Sparsification via Stochastic Binary Optimization
- FreezeNet: Full Performance by Reduced Storage Costs
- DASS: Differentiable Architecture Search for Sparse neural networks
- RIGA: Covert and Robust White-Box Watermarking of Deep Neural Networks
- Serpens: A High Bandwidth Memory Based Accelerator for General-Purpose Sparse Matrix-Vector Multiplication
- Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions
- Pruning artificial neural networks: a way to find well-generalizing, high-entropy sharp minima
- GeneCAI: Genetic Evolution for Acquiring Compact AI
- Taxonomy of Saliency Metrics for Channel Pruning
- Structured Deep Neural Network Pruning via Matrix Pivoting
- Optimizing for Interpretability in Deep Neural Networks with Tree Regularization
- A Thorough Performance Benchmarking on Lightweight Embedding-based Recommender Systems
- Sparseout: Controlling Sparsity in Deep Networks
- Efficient Hybrid Network Architectures for Extremely Quantized Neural Networks Enabling Intelligence at the Edge
- Build a Compact Binary Neural Network through Bit-level Sensitivity and Data Pruning
- Fast and Accurate, Convolutional Neural Network Based Approach for Object Detection from UAV
- Compressed Learning of Deep Neural Networks for OpenCL-Capable Embedded Systems
- A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural Networks
- Designing Adaptive Neural Networks for Energy-Constrained Image Classification
- HufuNet: Embedding the Left Piece as Watermark and Keeping the Right Piece for Ownership Verification in Deep Neural Networks
- EPNAS: Efficient Progressive Neural Architecture Search
- Edge-Cloud Collaborated Object Detection via Difficult-Case Discriminator
- Leveraging Sparse Linear Layers for Debuggable Deep Networks
- Block-wise Dynamic Sparseness
- Studying the Consistency and Composability of Lottery Ticket Pruning Masks
- CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay Compensation
- SpRRAM: A Predefined Sparsity Based Memristive Neuromorphic Circuit for Low Power Application
- SETGAN: Scale and Energy Trade-off GANs for Image Applications on Mobile Platforms
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
- A Quadratic Actor Network for Model-Free Reinforcement Learning
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- The staircase property: How hierarchical structure can guide deep learning
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Reliable Identification of Redundant Kernels for Convolutional Neural Network Compression
- Learning Connectivity of Neural Networks from a Topological Perspective
- Towards Design Methodology of Efficient Fast Algorithms for Accelerating Generative Adversarial Networks on FPGAs