1 citations · 1 across the 12 of their papers we have counts for
12 papers
Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization
Junjie Yao, Liangkai Hang, Zhi-Qin John Xu
Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern language models learn embeddings…
Metis: Memory Foundation Model
Zeyu Zhang, Ziliang Guo, Yihang Sun +14
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reaso…
Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay
Xiaolong Li, Zhangchen Zhou, Zhi-Qin John Xu
Most explanations of training instability focus on \emph{learning-rate criticality}, typically characterized by the Edge of Stability, beyond which optimization becomes unstable. W…
A First-Principles Theory of Slow Thinking and Active Perception
Hongkang Yang, Zhi-Qin John Xu, Feiyu Xiong +1
As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception. It formally derives s…
Small Initialization Matters for Large Language Models
Liangkai Hang, Junjie Yao, Zhiyu Li +3
Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although progress is usually attributed to…
Towards Understanding Adam Convergence on Highly Degenerate Polynomials
Zhiwei Bai, Jiajie Zhao, Zhangchen Zhou +2
Adam is a widely used optimization algorithm in deep learning, yet the specific class of objective functions where it exhibits inherent advantages remains underexplored. Unlike pri…