1 paper
Chenxuan Yao, Yuchong Hu, Feifan Liu +4
Distributed training of large deep-learning models often leads to failures, so checkpointing is commonly employed for recovery. State-of-the-art studies focus on frequent checkpoin…