collaborators

6 papers

cs.LG2026

Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1

In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer d…

cs.LG2026

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

Leena Chennuru Vankadara, Moritz Haas, Luke Hayward +2

Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hy…

cs.LG2025

Training Neural Networks at Any Scale

Thomas Pethick, Kimon Antonakopoulos, Antonio Silveti-Falls +2

This article reviews modern optimization methods for training neural networks with an emphasis on efficiency and scale. We present state-of-the-art optimization algorithms under a…

cs.LG2025

On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling

Moritz Haas, Sebastian Bordt, Ulrike von Luxburg +1

Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory…

cs.AI2025

The Amazon Nova Family of Models: Technical Report and Model Card

Amazon AGI, Aaron Langford, Aayush Shah +783

We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…

cs.LG2025

μP: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling

Moritz Haas, Jin Xu, Volkan Cevher +1

Sharpness Aware Minimization (SAM) enhances performance across various neural architectures and datasets. As models are continually scaled up to improve performance, a rigorous und…