Only Train Once: A One-Shot Neural Network Training And Pruning Framework
arXiv:2107.07467
Abstract
Structured pruning is a commonly used technique in deploying deep neural networks (DNNs) onto resource-constrained devices. However, the existing pruning methods are usually heuristic, task-specified, and require an extra fine-tuning procedure. To overcome these limitations, we propose a framework that compresses DNNs into slimmer architectures with competitive performances and significant FLOPs reductions by Only-Train-Once (OTO). OTO contains two keys: (i) we partition the parameters of DNNs into zero-invariant groups, enabling us to prune zero groups without affecting the output; and (ii) to promote zero groups, we then formulate a structured-sparsity optimization problem and propose a novel optimization algorithm, Half-Space Stochastic Projected Gradient (HSPG), to solve it, which outperforms the standard proximal methods on group sparsity exploration and maintains comparable convergence. To demonstrate the effectiveness of OTO, we train and compress full models simultaneously from scratch without fine-tuning for inference speedup and parameter reduction, and achieve state-of-the-art results on VGG16 for CIFAR10, ResNet50 for CIFAR10 and Bert for SQuAD and competitive result on ResNet50 for ImageNet. The source code is available at https://github.com/tianyic/only_train_once.
Accepted by NeurIPS 2021
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Nearly unbiased variable selection under minimax concave penalty
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- Learning Structured Sparsity in Deep Neural Networks
- The State of Sparsity in Deep Neural Networks
- Reducing Transformer Depth on Demand with Structured Dropout
- Comparing Rewinding and Fine-tuning in Neural Network Pruning
- Operation-Aware Soft Channel Pruning using Differentiable Masks
- Statistical Adaptive Stochastic Gradient Methods
- CDFI: Compression-Driven Network Design for Frame Interpolation
- Structured Sparsity Inducing Adaptive Optimizers for Deep Learning