7 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 6 cited
Horizontally Fused Training Array: An Effective Hardware Utilization Squeezer for Training Novel Deep Learning Models
Shang Wang, Peiming Yang, Yuxuan Zheng +2
Driven by the tremendous effort in researching novel deep learning (DL) algorithms, the training cost of developing new models increases staggeringly in recent years. We analyze GP…
cs.LG2020★ 7 cited
Multi-node Bert-pretraining: Cost-efficient Approach
Jiahuang Lin, Xin Li, Gennady Pekhimenko
Recently, large scale Transformer-based language models such as BERT, GPT-2, and XLNet have brought about exciting leaps in state-of-the-art results for many Natural Language Proce…