1 paper
Tenghui Ma, Jihu Guo, Wei Gao +4
Hybrid parallelism underpins large-scale LLM training across tens of thousands of GPUs. At such scale, hardware failures on individual devices lead to performance skew across devic…