3 papers
cs.DC2026
GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining
Jieying Wang, Shuyuan Fan, Mingkai Zheng +1
Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can signi…
cond-mat.stat-mech2025
Free energy dissipation and a decomposition of general jump diffusions on without detailed balance
Shuyuan Fan, Qi Zhang
We analyze the thermodynamic structure of jump diffusions combining Brownian and Poisson noise, a class of stochastic dynamics relevant to non-equilibrium statistical physics. For…
cs.DC2025
Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
Shuyuan Fan, Zhao Zhang
Global communication, such as all-reduce and allgather, is the prominent performance bottleneck in large language model (LLM) pretraining. To address this issue, we present Pier, a…