2 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2021★ 2 cited
Towards Structured Dynamic Sparse Pre-Training of BERT
Anastasia Dietrich, Frithjof Gressmann, Douglas Orr +3
Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a…
cs.CL2021★ 2 cited
GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
Ivan Chelombiev, Daniel Justus, Douglas Orr +4
Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requi…