8 citations · 18 across the 4 of their papers we have counts for
6 papers
Towards Structured Dynamic Sparse Pre-Training of BERT
Anastasia Dietrich, Frithjof Gressmann, Douglas Orr +3
Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a…
GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
Ivan Chelombiev, Daniel Justus, Douglas Orr +4
Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requi…
Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
Dominic Masters, Antoine Labatie, Zach Eaton-Rosen +1
Much recent research has been dedicated to improving the efficiency of training and inference for image classification. This effort has commonly focused on explicitly improving the…
Parallel Training of Deep Networks with Local Updates
Michael Laskin, Luke Metz, Seth Nabarro +5
Deep learning models trained on large data sets have been widely successful in both vision and language domains. As state-of-the-art deep learning architectures have continued to g…
Improving Neural Network Training in Low Dimensional Random Bases
Frithjof Gressmann, Zach Eaton-Rosen, Carlo Luschi
Stochastic Gradient Descent (SGD) has proven to be remarkably effective in optimizing deep neural networks that employ ever-larger numbers of parameters. Yet, improving the efficie…
Revisiting Small Batch Training for Deep Neural Networks
Dominic Masters, Carlo Luschi
Modern deep neural network training is typically based on mini-batch stochastic gradient optimization. While the use of large mini-batches increases the available computational par…