10 citations · 16 across the 4 of their papers we have counts for
3 papers · 1 filter
Slicing and Dicing: Configuring Optimal Mixtures of Experts
Margaret Li, Sneha Kudugunta, Danielle Rothermel +1
Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularity, shared experts, load balanc…
(Mis)Fitting: A Survey of Scaling Laws
Margaret Li, Sneha Kudugunta, Luke Zettlemoyer
Modern foundation models rely heavily on using scaling laws to guide crucial training decisions. Researchers often extrapolate the optimal architecture and hyper parameters setting…
A Loss Curvature Perspective on Training Instability in Deep Learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg +6
In this work, we study the evolution of the loss Hessian across many classification tasks in order to understand the effect the curvature of the loss has on the training dynamics.…