1 paper
Yanlin Long, Yufei Gu, Zeke Xie
Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Motivated by this challenge, w…