1 paper · 1 filter
Sibei Liu, Zhijian Hu
Learning rate (LR) schedules in large language model (LLM) training often follow empirical templates: warm-up, constant plateau/stable phase, and decay (WSD). However, the mechanis…