10 papers · 1 filter
Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
Yizhou Liu, Jeff Gore
Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power la…
Inverse Depth Scaling From Most Layers Being Similar
Yizhou Liu, Sara Kangaslahti, Ziming Liu +1
Neural scaling laws relate loss to model size in large language models (LLMs), yet depth and width may contribute to performance differently, requiring more detailed studies. Here,…
Universal One-third Time Scaling in Learning Peaked Distributions
Yizhou Liu, Ziming Liu, Cengiz Pehlevan +1
Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic a…
Superposition unifies power-law training dynamics
Zixin Jessie Chen, Hao Chen, Yizhou Liu +1
We investigate the role of feature superposition in the emergence of power-law training dynamics using a teacher-student framework. We first derive an analytic theory for training…
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
Qiyao Liang, Jinyeop Song, Yizhou Liu +4
Weight-perturbation evolution strategies (ES) can fine-tune billion-parameter language models with surprisingly small populations (e.g., ), contradicting classical…
Neural Thermodynamic Laws for Large Language Model Training
Ziming Liu, Yizhou Liu, Jeff Gore +1
Beyond neural scaling laws, little is known about the laws underlying large language models (LLMs). We introduce Neural Thermodynamic Laws (NTL) -- a new framework that offers fres…