Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
Jie Li, Chenxin Jia, Jinliang Shen +5
In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assumin…
cs.DC2026
X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference
Jianwen Xian, Zhiyuan Xu, Yuchen Li +8
Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Ten…