4 citations · 4 across the 1 of their papers we have counts for
3 papers
cs.LG2020★ 4 cited
CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery
Kiwan Maeng, Shivam Bharuka, Isabel Gao +8
The paper proposes and optimizes a partial recovery training system, CPR, for recommendation models. CPR relaxes the consistency requirement by enabling non-failed nodes to proceed…
cs.DC2020
Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems
Maxim Naumov, John Kim, Dheevatsa Mudigere +12
Large-scale training is important to ensure high performance and accuracy of machine-learning models. At Facebook we use many different models, including computer vision, video and…
cs.LG2020
ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training
Qinqing Zheng, Bor-Yiing Su, Jiyan Yang +7
Recommendation systems are often trained with a tremendous amount of data, and distributed training is the workhorse to shorten the training time. While the training throughput can…