2 papers
cs.DC2026
FlashBoot: Sub-Second Weight Loading for Large Models at Rack Scale
Issac Zhu, Hscos Zhang, Keith Jiang +3
Flagship Mixture-of-Experts (MoE) models are growing fast along two axes at once: total parameter count and the number of experts. In elastic deployment scenarios, many GPUs across…
cs.LG2026
SparseRL-Sync: Lossless Weight Synchronization with ~100x Less Communication
Lucas Hu, Ranchi Zhao, Isaac Zhu +4
In large-scale reinforcement learning (RL) systems with decoupled Trainer-Rollout execution, the Trainer must regularly synchronize policy weights to the Rollout side to limit poli…