1 paper
Shuchen Zhu, Yuxin Fang, Mingze Wang +1
Pretraining accounts for a large fraction of the total computational cost in LLM training. However, noise-dominant gradients and the highly ill-conditioned loss landscape bring sev…