3 citations · 7 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2021★ 2 cited
Towards Structured Dynamic Sparse Pre-Training of BERT
Anastasia Dietrich, Frithjof Gressmann, Douglas Orr +3
Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a…
cs.CL2021★ 2 cited
GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
Ivan Chelombiev, Daniel Justus, Douglas Orr +4
Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requi…