Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
CIERA: Cross-Iteration Exponent Reuse for Lossless Allgather in Sharded MoE Training
Ali Zafar Sadiq, Haiying Shen, Masahiro Tanaka
In training Mixture-of-Experts (MoE) models, sharded data parallelism partitions each expert's parameters across GPUs, requiring an Allgather operation to reconstruct the full weig…
cs.DC2024
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
Haiying Shen, Tanmoy Sen
As Large Language Models (LLMs) continue to grow, reducing costs and alleviating GPU demands has become increasingly critical. However, existing schedulers primarily target either…