2 papers
cs.LG2026
Superposition unifies power-law training dynamics
Zixin Jessie Chen, Hao Chen, Yizhou Liu +1
We investigate the role of feature superposition in the emergence of power-law training dynamics using a teacher-student framework. We first derive an analytic theory for training…
cs.LG2025
Neural Thermodynamic Laws for Large Language Model Training
Ziming Liu, Yizhou Liu, Jeff Gore +1
Beyond neural scaling laws, little is known about the laws underlying large language models (LLMs). We introduce Neural Thermodynamic Laws (NTL) -- a new framework that offers fres…