1 paper
Xueze Kang, Guangyu Xiang, Yuxin Wang +16
Large-scale LLM pretraining now runs across 105--106 accelerators, making failures routine and elasticity mandatory. We posit that an elastic-native training system must join…