2 papers
cs.LG2026
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
Wuyue Zhang, Chongdong Huang, Chunbo You +3
Training large-scale Mixture-of-Experts (MoE) models is bottlenecked by activation memory and expert-parallel communication, yet FP4 training remains impractical on Hopper-class GP…
cs.LG2025
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
Fengjuan Wang, Zhiyi Su, Xingzhu Hu +2
Training large Mixture-of-Experts (MoE) models remains computationally prohibitive due to their extreme compute and memory demands. Although low-precision training promises to acce…