1 paper
Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2
AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying…