The State of Sparsity in Deep Neural Networks
arXiv:1902.09574
Abstract
We rigorously evaluate three state-of-the-art techniques for inducing sparsity in deep neural networks on two large-scale learning tasks: Transformer trained on WMT 2014 English-to-German, and ResNet-50 trained on ImageNet. Across thousands of experiments, we demonstrate that complex techniques (Molchanov et al., 2017; Louizos et al., 2017b) shown to yield high compression rates on smaller datasets perform inconsistently, and that simple magnitude pruning approaches achieve comparable or better results. Additionally, we replicate the experiments performed by (Frankle & Carbin, 2018) and (Liu et al., 2018) at scale and show that unstructured sparse architectures learned through pruning cannot be trained from scratch to the same test set performance as a model trained with joint sparsification and optimization. Together, these results highlight the need for large-scale benchmarks in the field of model compression. We open-source our code, top performing model checkpoints, and results of all hyperparameter configurations to establish rigorous baselines for future work on compression and sparsification.
References in corpus (2)
Cited by in corpus (17)
- Continual Learning via Neural Pruning
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Looking GLAMORous: Vehicle Re-Id in Heterogeneous Cameras Networks with Global and Local Attention
- Lookahead: A Far-Sighted Alternative of Magnitude-based Pruning
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation
- CAT: Compression-Aware Training for bandwidth reduction
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Privacy-preserving Learning via Deep Net Pruning
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- word2ket: Space-efficient Word Embeddings inspired by Quantum Entanglement
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- On Iterative Neural Network Pruning, Reinitialization, and the Similarity of Masks
- Small, Accurate, and Fast Vehicle Re-ID on the Edge: the SAFR Approach
- Dissecting Pruned Neural Networks
- Grassmannian Packings in Neural Networks: Learning with Maximal Subspace Packings for Diversity and Anti-Sparsity