collaborators

6 papers

cs.CL2026

Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers

Zixuan Gong, Shijia Li, Yong Liu +1

Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from…

cs.LG2026

Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry

Ye Su, Huayi Tang, Zixuan Gong +1

While Mixture-of-Experts (MoE) architectures define the state-of-the-art, their theoretical success is often attributed to heuristic efficiency rather than geometric expressivity.…

cs.CL2026

Beyond the Black Box: A Survey on the Theory and Mechanism of Large Language Models

Zeyu Gan, Ruifeng Ren, Wei Yao +9

The rapid emergence of Large Language Models (LLMs) has precipitated a profound paradigm shift in Artificial Intelligence, delivering monumental engineering successes that increasi…

cs.LG2026

Effective Frontiers: A Unification of Neural Scaling Laws

Jiaxuan Zou, Zixuan Gong, Ye Su +2

Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity (), datasize (), and compute (). However, existing theoretical…

cs.LG2026

What Makes Looped Transformers Perform Better Than Non-Recursive Ones

Zixuan Gong, Yong Liu, Jiaye Teng

While looped transformers (termed as Looped-Attn) often outperform standard transformers (termed as Single-Attn) on complex reasoning tasks, the mechanism for this advantage remain…

cs.CL2025

Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization

Zixuan Gong, Xiaolin Hu, Huayi Tang +1

Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: (a)…