1 paper
Daemyung Kang, Eunjin Hwang, Hanjeong Lee +11
Large-scale AI training is fundamentally a distributed systems problem, where hardware failures are routine operating conditions rather than rare exceptions, yet public operational…