5 papers
Taking the Road Less Scheduled with Adaptive Polyak Steps
Dimitris Oikonomou, Matthew Buchholz, Yuen-Man Pun +2
Schedule-Free SGD, proposed in [Defazio et al., 2024], achieves optimal convergence rates without requiring the training horizon in advance, by replacing learning rate schedules wi…
Fisher meets Feynman: score-based variational inference with a product of experts
Diana Cai, Robert M. Gower, David M. Blei +1
We introduce a highly expressive yet distinctly tractable family for black-box variational inference (BBVI). Each member of this family is a weighted product of experts (PoE), and…
An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
Michael Crawshaw, Chirag Modi, Mingrui Liu +1
To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization.…
Level Set Teleportation: An Optimization Perspective
Aaron Mishkin, Alberto Bietti, Robert M. Gower
We study level set teleportation, an optimization routine which tries to accelerate gradient descent (GD) by maximizing the gradient norm over a level set of the objective. While t…
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
Aaron Mishkin, Ahmed Khaled, Yuanhao Wang +2
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case con…