Search Spaces for Neural Model Training
arXiv:2105.12920
Abstract
While larger neural models are pushing the boundaries of what deep learning can do, often more weights are needed to train models rather than to run inference for tasks. This paper seeks to understand this behavior using search spaces -- adding weights creates extra degrees of freedom that form new paths for optimization (or wider search spaces) rendering neural model training more effective. We then show how we can augment search spaces to train sparse models attaining competitive scores across dozens of deep learning workloads. They are also are tolerant of structures targeting current hardware, opening avenues for training and inference acceleration. Our work encourages research to explore beyond massive neural models being used today.
References in corpus (15)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Scaling Laws for Neural Language Models
- The Loss Surfaces of Multilayer Networks
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- The State of Sparsity in Deep Neural Networks
- Deep Learning Scaling is Predictable, Empirically
- Comparing Rewinding and Fine-tuning in Neural Network Pruning
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Scaling Laws for Autoregressive Generative Modeling
- Block-Sparse Recurrent Neural Networks
- Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Top-KAST: Top-K Always Sparse Training
- Accelerating Sparse Deep Neural Networks