2 citations · 2 across the 1 of their papers we have counts for
1 paper
Yuchen Zhong, Guangming Sheng, Juncheng Liu +2
As the size of deep learning models gets larger and larger, training takes longer time and more resources, making fault tolerance more and more critical. Existing state-of-the-art…