1 paper
Shujie Han, Feng Jiang, Patrick P. C. Lee +5
Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster-wide outages. Existing checkp…