2 papers
cs.DC2026
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
Yipeng Liu, Chang Liu, Si Shen +16
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges bey…
cs.PF2021
Pinpointing the Memory Behaviors of DNN Training
Jiansong Li, Xiao Dong, Guangli Li +9
The training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN train…