6 citations · 13 across the 14 of their papers we have counts for
1 paper · 2 filters
Siyuan Li, Juanxi Tian, Zedong Wang +4
Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variat…