activity
20182022
most citedShadowing Properties of Optimization Algorithms

10 citations · 16 across the 8 of their papers we have counts for

collaborators

13 papers

math.OC20221 cited

On the Theoretical Properties of Noise Correlation in Stochastic Optimization

Aurelien Lucchi, Frank Proske, Antonio Orvieto +2

Studying the properties of stochastic noise to optimize complex non-convex functions has been an active area of research in the field of machine learning. Prior work has shown that…

math.OC2021

On the Second-order Convergence Properties of Random Search Methods

Aurelien Lucchi, Antonio Orvieto, Adamos Solomou

We study the theoretical convergence properties of random-search methods when optimizing non-convex objective functions without having access to derivatives. We prove that standard…

math.OC2021

Rethinking the Variational Interpretation of Nesterov's Accelerated Method

Peiyuan Zhang, Antonio Orvieto, Hadi Daneshmand

The continuous-time model of Nesterov's momentum provides a thought-provoking perspective for understanding the nature of the acceleration phenomenon in convex optimization. One of…

cs.LG2021

Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks

Antonio Orvieto, Jonas Kohler, Dario Pavllo +2

This paper revisits the so-called vanishing gradient phenomenon, which commonly occurs in deep randomly initialized neural networks. Leveraging an in-depth analysis of neural chain…

math.OC2021

Revisiting the Role of Euler Numerical Integration on Acceleration and Stability in Convex Optimization

Peiyuan Zhang, Antonio Orvieto, Hadi Daneshmand +2

Viewing optimization methods as numerical integrators for ordinary differential equations (ODEs) provides a thought-provoking modern framework for studying accelerated first-order…

cs.LG20201 cited

Two-Level K-FAC Preconditioning for Deep Learning

Nikolaos Tselepidis, Jonas Kohler, Antonio Orvieto

In the context of deep learning, many optimization methods use gradient covariance information in order to accelerate the convergence of Stochastic Gradient Descent. In particular,…