1 citations · 1 across the 17 of their papers we have counts for
24 papers
How the Hessian-Spectrum of Neural Networks Depends on Data
Jasraj Singh, Enea Monzio Compagnoni, Antonio Orvieto
The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of ge…
Beyond a Single Explanation of the Adam--SGD Gap
Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3
Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…
GRASP: Deterministic argument ranking in interaction graphs
Diganta Misra, Antonio Orvieto, Rediet Abebe +1
Large language models are increasingly deployed as automated judges to evaluate the strength of arguments. As this role expands, their legitimacy depends on consistency, transparen…
Muown: Row-Norm Control for Muon Optimization
Kai Lion, Florian Hübler, Bingcong Li +2
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon wi…
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
Egor Shulgin, Dimitri von Rütte, Tianyue H. Zhang +3
Hyperparameter transfer has become an important component of modern large-scale training recipes. Existing methods, such as muP, primarily focus on transfer between model sizes, wi…
GASP: Guided Asymmetric Self-Play For Coding LLMs
Swadesh Jana, Cansu Sancaktar, Tomáš Daniš +3
Asymmetric self-play has emerged as a promising paradigm for post-training large language models, where a teacher continually generates questions for a student to solve at the edge…