4 citations · 4 across the 2 of their papers we have counts for
3 papers
cs.PF2026
Compass: Dissecting Communication and Computation Operators for Efficient LLM Training
Guangyu Xiang, Lin Zhang, Haoxuan Yu +3
Overlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existi…
cs.DC2024★ 4 cited
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
Zichen Tang, Junlin Huang, Rudan Yan +5
Current data compression methods, such as sparsification in Federated Averaging (FedAvg), effectively enhance the communication efficiency of Federated Learning (FL). However, thes…
cs.DC2024
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
Xinglin Pan, Wenxiang Lin, Shaohuai Shi +3
Sparsely-activated Mixture-of-Expert (MoE) layers have found practical applications in enlarging the model size of large-scale foundation models, with only a sub-linear increase in…