Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding
Ming Gao, Yanwu Xu, Hao Zhang
Matrix optimizers such as Muon are attractive for large-scale training because they can improve convergence and token efficiency over coordinate-wise optimizers. Muon does this by…
cs.DC2024
Training Overhead Ratio: A Practical Reliability Metric for Large Language Model Training Systems
Ning Lu, Qian Xie, Hao Zhang +4
Large Language Models (LLMs) are revolutionizing the AI industry with their superior capabilities. Training these models requires large-scale GPU clusters and significant computing…