4 papers
Dimension-Free Saddle-Point Escape in Muon
Yanlin Long, Yufei Gu, Zeke Xie
Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Motivated by this challenge, w…
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
Yulong Huang, Xiang Liu, Hongxiang Huang +5
Linear Attention (LA) offers a promising paradigm for scaling large language models (LLMs) to long sequences by avoiding the quadratic complexity of self-attention. Recent LA model…
Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
Ji Zhao, Yufei Gu, Shitong Shao +3
As Large Language Models (LLMs) achieve remarkable empirical success through scaling model and data size, pretraining has become increasingly critical yet computationally prohibiti…
Mano: Restriking Manifold Optimization for LLM Training
Yufei Gu, Zeke Xie
While large language models (LLMs) have emerged as a significant advancement in artificial intelligence, the hardware and computational costs for training LLMs are also significant…