2 papers
cs.LG2026
Dissecting Outlier Dynamics in LLM NVFP4 Pretraining
Peijie Dong, Ruibo Fan, Yuechen Tao +11
Training large language models using 4-bit arithmetic enhances throughput and memory efficiency. Yet, the limited dynamic range of FP4 increases sensitivity to outliers. While NVFP…
cs.DC2025
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Wentao Liu, Yuhao Hu, Ruiting Zhou +2
Mixture-of-Experts (MoE) has become a dominant architecture in large language models (LLMs) due to its ability to scale model capacity via sparse expert activation. Meanwhile, serv…