activity
20182021
most citedMaking EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training

8 citations · 18 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL20212 cited

Towards Structured Dynamic Sparse Pre-Training of BERT

Anastasia Dietrich, Frithjof Gressmann, Douglas Orr +3

Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a…

cs.CL20212 cited

GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures

Ivan Chelombiev, Daniel Justus, Douglas Orr +4

Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requi…

cs.LG20218 cited

Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training

Dominic Masters, Antoine Labatie, Zach Eaton-Rosen +1

Much recent research has been dedicated to improving the efficiency of training and inference for image classification. This effort has commonly focused on explicitly improving the…

cs.LG2020

Parallel Training of Deep Networks with Local Updates

Michael Laskin, Luke Metz, Seth Nabarro +5

Deep learning models trained on large data sets have been widely successful in both vision and language domains. As state-of-the-art deep learning architectures have continued to g…

cs.LG20206 cited

Improving Neural Network Training in Low Dimensional Random Bases

Frithjof Gressmann, Zach Eaton-Rosen, Carlo Luschi

Stochastic Gradient Descent (SGD) has proven to be remarkably effective in optimizing deep neural networks that employ ever-larger numbers of parameters. Yet, improving the efficie…

cs.LG2018

Revisiting Small Batch Training for Deep Neural Networks

Dominic Masters, Carlo Luschi

Modern deep neural network training is typically based on mini-batch stochastic gradient optimization. While the use of large mini-batches increases the available computational par…