3 papers
cs.LG2026
Neural Networks Provably Learn Spectral Representations for Group Composition
Jianliang He, Leda Wang, Fengzhuo Zhang +2
Understanding how structured internal structure emerges during neural network training is central to the study of deep learning. We investigate this phenomenon through the group co…
cs.LG2026
On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking
Jianliang He, Leda Wang, Siyu Chen +1
We present a comprehensive analysis of how two-layer neural networks learn features to solve the modular addition task. Our work provides a full mechanistic interpretation of the l…
cs.LG2025
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
Jianliang He, Xintian Pan, Siyu Chen +1
We study how multi-head softmax attention models are trained to perform in-context learning on linear data. Through extensive empirical experiments and rigorous theoretical analysi…