37 citations · 37 across the 3 of their papers we have counts for
3 papers
cs.DC2022
Efficient AlphaFold2 Training using Parallel Evoformer and Branch Parallelism
Guoxia Wang, Zhihua Wu, Xiaomin Fang +4
The accuracy of AlphaFold2, a frontier end-to-end structure prediction system, is already close to that of the experimental determination techniques. Due to the complex model archi…
cs.CL2021★ 37 cited
ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
Shuohuan Wang, Yu Sun, Yang Xiang +26
Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models c…
cs.DC2021
End-to-end Adaptive Distributed Training on PaddlePaddle
Yulong Ao, Zhihua Wu, Dianhai Yu +10
Distributed training has become a pervasive and effective approach for training a large neural network (NN) model with processing massive data. However, it is very challenging to s…