2 citations · 4 across the 9 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao +1
Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every pea…
cs.DC2026
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao +3
This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes various existing and novel paral…