6 papers
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
Zixuan Gong, Shijia Li, Yong Liu +1
Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from…
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
Ye Su, Huayi Tang, Zixuan Gong +1
While Mixture-of-Experts (MoE) architectures define the state-of-the-art, their theoretical success is often attributed to heuristic efficiency rather than geometric expressivity.…
Beyond the Black Box: A Survey on the Theory and Mechanism of Large Language Models
Zeyu Gan, Ruifeng Ren, Wei Yao +9
The rapid emergence of Large Language Models (LLMs) has precipitated a profound paradigm shift in Artificial Intelligence, delivering monumental engineering successes that increasi…
Effective Frontiers: A Unification of Neural Scaling Laws
Jiaxuan Zou, Zixuan Gong, Ye Su +2
Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity (), datasize (), and compute (). However, existing theoretical…
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
Zixuan Gong, Yong Liu, Jiaye Teng
While looped transformers (termed as Looped-Attn) often outperform standard transformers (termed as Single-Attn) on complex reasoning tasks, the mechanism for this advantage remain…
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
Zixuan Gong, Xiaolin Hu, Huayi Tang +1
Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: (a)…