8 papers
AdaGrad does not adapt to Hölder-smoothness for composite objectives
Matia Bojovic, Saverio Salzo, Massimiliano Pontil
We exhibit a simple deterministic one-dimensional convex composite optimization problem for which AdaGrad scheme does not achieve the classical convergence rate $\mathcal{O}(n^{-(1…
Bilevel learning
Riccardo Grazzi, Massimiliano Pontil, Saverio Salzo +1
Bilevel learning refers to machine learning problems that can be formulated as bilevel optimization models, where decisions are organized in a hierarchical structure. This paradigm…
Iteration Complexity of Frank-Wolfe and Its Variants for Bilevel Optimization
Anthony Palmieri, Francesco Rinaldi, Saverio Salzo +1
We study Frank-Wolfe (FW) methods for constrained bilevel optimization when the lower-level problem is solved only approximately, yielding biased and inexact hypergradients. We ana…
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
Matia Bojovic, Saverio Salzo, Massimiliano Pontil
Vanilla gradient methods are often highly sensitive to the choice of stepsize, which typically requires manual tuning. Adaptive methods alleviate this issue and have therefore beco…
The iterates of FISTA converge even under inexact computations and stochastic gradients
Saverio Salzo
Very recently, the papers "Point Convergence of Nesterov's Accelerated Gradient Method: An AI-Assisted Proof" by Jang and Ryu, and "The Iterates of Nesterov's Accelerated Algorithm…
Convergence Properties of Stochastic Hypergradients
Riccardo Grazzi, Massimiliano Pontil, Saverio Salzo
Bilevel optimization problems are receiving increasing attention in machine learning as they provide a natural framework for hyperparameter optimization and meta-learning. A key st…