Towards Structured Dynamic Sparse Pre-Training of BERT
arXiv:2108.06277
Abstract
Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a straightforward, dynamic always-sparse pre-training approach for BERT language modeling task, which leverages periodic compression steps based on magnitude pruning followed by random parameter re-allocation. This approach enables us to achieve Pareto improvements in terms of the number of floating-point operations (FLOPs) over statically sparse and dense models across a broad spectrum of network sizes. Furthermore, we demonstrate that training remains FLOP-efficient when using coarse-grained block sparsity, making it particularly promising for efficient execution on modern hardware accelerators.
References in corpus (10)
- Scaling Laws for Neural Language Models
- The State of Sparsity in Deep Neural Networks
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Carbon Emissions and Large Neural Network Training
- Block-Sparse Recurrent Neural Networks
- Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch
- BASE Layers: Simplifying Training of Large, Sparse Models
- Top-KAST: Top-K Always Sparse Training