3 papers
cs.LG2026
The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity
Siquan Li, Kaiqi Jiang, Jiacheng Sun +1
Despite the prevalence of the attention sink phenomenon in Large Language Models (LLMs), where initial tokens disproportionately monopolize attention scores, its structural origins…
cs.LG2025
Understanding the Evolution of the Neural Tangent Kernel at the Edge of Stability
Kaiqi Jiang, Jeremy Cohen, Yuanzhi Li
The study of Neural Tangent Kernels (NTKs) in deep learning has drawn increasing attention in recent years. NTKs typically actively change during training and are related to featur…
cs.LG2025
Fairness Risks for Group-conditionally Missing Demographics
Kaiqi Jiang, Wenzhe Fan, Mao Li +1
Fairness-aware classification models have gained increasing attention in recent years as concerns grow on discrimination against some demographic groups. Most existing models requi…