1 paper
Xianzhe Meng, Qiangsheng Zeng, Ling Luo +7
Training stability is typically regarded as a prerequisite for reliable optimization in large language models. In this work, we analyze how stabilizing training dynamics affects th…