2 papers
cs.LG2026
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint
Haotian Xie, Junlin Chen, Mingkai Zheng +2
State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware…
cs.LG2026
SCAPE: Accurate and Efficient LLM Training with Extreme Sparse Communication
Mingkai Zheng, Junlin Chen, Haotian Xie +1
Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronizatio…