4 papers
Dimension-Free Saddle-Point Escape in Muon
Yanlin Long, Yufei Gu, Zeke Xie
Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Motivated by this challenge, w…
FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters
Shitong Shao, Yufei Gu, Zeke Xie
The recent advent of powerful video generation models, such as Hunyuan, WanX, Veo3, and Kling, has inaugurated a new era in the field. However, the practical deployment of these mo…
Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
Ji Zhao, Yufei Gu, Shitong Shao +3
As Large Language Models (LLMs) achieve remarkable empirical success through scaling model and data size, pretraining has become increasingly critical yet computationally prohibiti…
Mano: Restriking Manifold Optimization for LLM Training
Yufei Gu, Zeke Xie
While large language models (LLMs) have emerged as a significant advancement in artificial intelligence, the hardware and computational costs for training LLMs are also significant…