6 citations · 10 across the 3 of their papers we have counts for
3 papers
Towards Structured Dynamic Sparse Pre-Training of BERT
Anastasia Dietrich, Frithjof Gressmann, Douglas Orr +3
Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a…
GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
Ivan Chelombiev, Daniel Justus, Douglas Orr +4
Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requi…
Improving Neural Network Training in Low Dimensional Random Bases
Frithjof Gressmann, Zach Eaton-Rosen, Carlo Luschi
Stochastic Gradient Descent (SGD) has proven to be remarkably effective in optimizing deep neural networks that employ ever-larger numbers of parameters. Yet, improving the efficie…