2 papers
cs.DC2026
ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving
Haipeng Yuan, Kaining Zheng, Yongshu Bai +5
Current large language model (LLM) inference systems universally deploy ultra-large-scale models using a combination of Tensor Parallelism (TP) and Pipeline Parallelism (PP). Howev…
cs.LG2026
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
Yuanqing Wang, Yuchen Zhang, Hao Lin +9
Modern large language model (LLM) training is inherently dynamic: resource fluctuations, RLHF phase shifts, and cluster elasticity continually reshape the optimal parallelism layou…