6 papers
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1
In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer d…
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
Leena Chennuru Vankadara, Moritz Haas, Luke Hayward +2
Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hy…
Training Neural Networks at Any Scale
Thomas Pethick, Kimon Antonakopoulos, Antonio Silveti-Falls +2
This article reviews modern optimization methods for training neural networks with an emphasis on efficiency and scale. We present state-of-the-art optimization algorithms under a…
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling
Moritz Haas, Sebastian Bordt, Ulrike von Luxburg +1
Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
μP: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
Moritz Haas, Jin Xu, Volkan Cevher +1
Sharpness Aware Minimization (SAM) enhances performance across various neural architectures and datasets. As models are continually scaled up to improve performance, a rigorous und…