2 papers
cs.DC2026
CIERA: Cross-Iteration Exponent Reuse for Lossless Allgather in Sharded MoE Training
Ali Zafar Sadiq, Haiying Shen, Masahiro Tanaka
In training Mixture-of-Experts (MoE) models, sharded data parallelism partitions each expert's parameters across GPUs, requiring an Allgather operation to reconstruct the full weig…
cs.OS2026
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
Shutian Luo, Ali Zafar Sadiq, Rui Yang +4
Modern LLM serving is increasingly serverless in shape: large model catalogs, long-tail invocations, and multi-tenant demand. Existing GPU serving systems face a tradeoff: dedicate…