7 papers · 1 filter
AdaGrad does not adapt to Hölder-smoothness for composite objectives
Matia Bojovic, Saverio Salzo, Massimiliano Pontil
We exhibit a simple deterministic one-dimensional convex composite optimization problem for which AdaGrad scheme does not achieve the classical convergence rate $\mathcal{O}(n^{-(1…
Bilevel learning
Riccardo Grazzi, Massimiliano Pontil, Saverio Salzo +1
Bilevel learning refers to machine learning problems that can be formulated as bilevel optimization models, where decisions are organized in a hierarchical structure. This paradigm…
Iteration Complexity of Frank-Wolfe and Its Variants for Bilevel Optimization
Anthony Palmieri, Francesco Rinaldi, Saverio Salzo +1
We study Frank-Wolfe (FW) methods for constrained bilevel optimization when the lower-level problem is solved only approximately, yielding biased and inexact hypergradients. We ana…
The iterates of FISTA converge even under inexact computations and stochastic gradients
Saverio Salzo
Very recently, the papers "Point Convergence of Nesterov's Accelerated Gradient Method: An AI-Assisted Proof" by Jang and Ryu, and "The Iterates of Nesterov's Accelerated Algorithm…
An Improved Analysis of the Clipped Stochastic subGradient Method under Heavy-Tailed Noise
Daniela Angela Parletta, Andrea Paudice, Saverio Salzo
In this paper, we provide novel optimal (or near optimal) convergence rates for a clipped version of the stochastic subgradient method. We consider nonsmooth convex problems over p…
Variance reduction techniques for stochastic proximal point algorithms
Cheik Traoré, Vassilis Apidopoulos, Saverio Salzo +1
In the context of finite sums minimization, variance reduction techniques are widely used to improve the performance of state-of-the-art stochastic gradient methods. Their practica…