2 papers
cs.DC2026
ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving
Haipeng Yuan, Kaining Zheng, Yongshu Bai +5
Current large language model (LLM) inference systems universally deploy ultra-large-scale models using a combination of Tensor Parallelism (TP) and Pipeline Parallelism (PP). Howev…
cs.LG2025
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
Yanxin Peng, Qingping Li, Baodong Wu +4
As large language models (LLMs) continue to grow in size and complexity, efficient checkpoint saving\&loading has become crucial for managing storage, memory usage, and fault toler…