12 citations · 14 across the 2 of their papers we have counts for
2 papers
cs.CL2021★ 2 cited
GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
Ivan Chelombiev, Daniel Justus, Douglas Orr +4
Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requi…
cs.DC2019★ 12 cited
CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
Alexandros Koliousis, Pijika Watcharapichat, Matthias Weidlich +3
Deep learning models are trained on servers with many GPUs, and training must scale with the number of GPUs. Systems such as TensorFlow and Caffe2 train models with parallel synchr…