13 citations · 13 across the 2 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Scalable Training of Mixture-of-Experts Models with Megatron Core
Zijie Yan, Hongxiao Bai, Xin Yao +42
Scaling Mixture-of-Experts (MoE) training introduces systems challenges absent in dense models. Because each token activates only a subset of experts, this sparsity allows total pa…
cs.DC2020★ 13 cited
Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity
Cong Guo, Bo Yang Hsueh, Jingwen Leng +7
Network pruning can reduce the high computation cost of deep neural network (DNN) models. However, to maintain their accuracies, sparse models often carry randomly-distributed weig…