43 citations · 114 across the 40 of their papers we have counts for
4 papers · 1 filter
Two-Level K-FAC Preconditioning for Deep Learning
Nikolaos Tselepidis, Jonas Kohler, Antonio Orvieto
In the context of deep learning, many optimization methods use gradient covariance information in order to accelerate the convergence of Stochastic Gradient Descent. In particular,…
Learning explanations that are hard to vary
Giambattista Parascandolo, Alexander Neitz, Antonio Orvieto +2
In this paper, we investigate the principle that `good explanations are hard to vary' in the context of deep learning. We show that averaging gradients across examples -- akin to a…
An Accelerated DFO Algorithm for Finite-sum Convex Functions
Yuwen Chen, Antonio Orvieto, Aurelien Lucchi
Derivative-free optimization (DFO) has recently gained a lot of momentum in machine learning, spawning interest in the community to design faster methods for problems where gradien…
Momentum Improves Optimization on Riemannian Manifolds
Foivos Alimisis, Antonio Orvieto, Gary Bécigneul +1
We develop a new Riemannian descent algorithm that relies on momentum to improve over existing first-order methods for geodesically convex optimization. In contrast, accelerated co…