2 papers
cs.LG2025
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
Shikai Qiu, Lechao Xiao, Andrew Gordon Wilson +2
What scaling limits govern neural network training dynamics when model size and training time grow in tandem? We show that despite the complex interactions between architecture, tr…
cs.LG2025
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
Atish Agarwala, Jeffrey Pennington
Recent empirical and theoretical work has shown that the dynamics of the large eigenvalues of the training loss Hessian have some remarkably robust features across models and datas…