collaborators

12 papers

cs.LG2026

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

Yizhou Liu, Jeff Gore

Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power la…

cs.LG2026

Inverse Depth Scaling From Most Layers Being Similar

Yizhou Liu, Sara Kangaslahti, Ziming Liu +1

Neural scaling laws relate loss to model size in large language models (LLMs), yet depth and width may contribute to performance differently, requiring more detailed studies. Here,…

cs.LG2026

Universal One-third Time Scaling in Learning Peaked Distributions

Yizhou Liu, Ziming Liu, Cengiz Pehlevan +1

Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic a…

cs.AI2026

Estimating the Empowerment of Language Model Agents

Jinyeop Song, Jeff Gore, Max Kleiman-Weiner

As language model (LM) agents become increasingly capable and adopted in real-world applications, there is a growing need for scalable evaluation frameworks beyond costly, manually…

cs.CV2026

Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution

Zixin Jessie Chen, Zhuo Chen, Archer Wang +4

Creating images from noise is image generation; reconstructing fine details from coarse inputs is super-resolution. Despite their practical differences, both can be understood as r…

cs.LG2026

Superposition unifies power-law training dynamics

Zixin Jessie Chen, Hao Chen, Yizhou Liu +1

We investigate the role of feature superposition in the emergence of power-law training dynamics using a teacher-student framework. We first derive an analytic theory for training…