3 papers
cs.DC2026
MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding
Ming Gao, Yanwu Xu, Hao Zhang
Matrix optimizers such as Muon are attractive for large-scale training because they can improve convergence and token efficiency over coordinate-wise optimizers. Muon does this by…
cs.LG2026
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
Zhichen Liu, Tianle Lun, Zhibin Wen +7
The paradigm of scaling Large Language Models (LLMs) in both parameter size and test time has pushed the boundaries of AI capabilities, but at the cost of making the traditional ge…
cs.DC2024
Training Overhead Ratio: A Practical Reliability Metric for Large Language Model Training Systems
Ning Lu, Qian Xie, Hao Zhang +4
Large Language Models (LLMs) are revolutionizing the AI industry with their superior capabilities. Training these models requires large-scale GPU clusters and significant computing…