20 citations · 20 across the 1 of their papers we have counts for
2 papers
cs.DC2021★ 20 cited
Maximizing Parallelism in Distributed Training for Huge Neural Networks
Zhengda Bian, Qifan Xu, Boxiang Wang +1
The recent Natural Language Processing techniques have been refreshing the state-of-the-art performance at an incredible speed. Training huge language models is therefore an impera…
cs.LG2021
An Efficient 2D Method for Training Super-Large Deep Learning Models
Qifan Xu, Shenggui Li, Chaoyu Gong +1
Huge neural network models have shown unprecedented performance in real-world applications. However, due to memory constraints, model parallelism must be utilized to host large mod…