4 papers
Kaczmarz Linear Attention
Jiaxuan Zou, Ruifeng Ren, Yong Liu
Long-context language modeling remains central to modern sequence modeling, but the quadratic cost of Transformer attention makes scaling computationally prohibitive. Linear recurr…
Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer
Jinghui Yuan, Jiaxuan Zou, Shuo Wang +2
Matrix-based optimizers have demonstrated immense potential in training Large Language Models (LLMs), however, designing an ideal optimizer remains a formidable challenge. A superi…
Effective Frontiers: A Unification of Neural Scaling Laws
Jiaxuan Zou, Zixuan Gong, Ye Su +2
Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity (), datasize (), and compute (). However, existing theoretical…
Capabilities and Fundamental Limits of Latent Chain-of-Thought
Jiaxuan Zou, Yaozhong Xiong, Yong Liu
Latent Chain-of-Thought (Latent CoT) models promise efficient reasoning via continuous representations, yet exhibit puzzling performance inconsistencies: excelling at exploration (…