1 paper
Wuyue Zhang, Chongdong Huang, Chunbo You +3
Training large-scale Mixture-of-Experts (MoE) models is bottlenecked by activation memory and expert-parallel communication, yet FP4 training remains impractical on Hopper-class GP…