4 citations · 9 across the 9 of their papers we have counts for
5 papers · 1 filter
PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient Estimation
Matilde Gargiani, Andrea Zanelli, Andrea Martinelli +2
Despite their success, policy gradient methods suffer from high variance of the gradient estimate, which can result in unsatisfactory sample complexity. Recently, numerous variance…
Convergence Analysis of Homotopy-SGD for non-convex optimization
Matilde Gargiani, Andrea Zanelli, Quoc Tran-Dinh +2
First-order stochastic methods for solving large-scale non-convex optimization problems are widely used in many big-data applications, e.g. training deep neural networks as well as…
On the Promise of the Stochastic Generalized Gauss-Newton Method for Training DNNs
Matilde Gargiani, Andrea Zanelli, Moritz Diehl +1
Following early work on Hessian-free methods for deep learning, we study a stochastic generalized Gauss-Newton method (SGN) for training DNNs. SGN is a second-order optimization me…
Probabilistic Rollouts for Learning Curve Extrapolation Across Hyperparameter Settings
Matilde Gargiani, Aaron Klein, Stefan Falkner +1
We propose probabilistic models that can extrapolate learning curves of iterative machine learning algorithms, such as stochastic gradient descent for training deep networks, based…
A Distributed Second-Order Algorithm You Can Trust
Celestine Dünner, Aurelien Lucchi, Matilde Gargiani +3
Due to the rapid growth of data and computational resources, distributed optimization has become an active research area in recent years. While first-order methods seem to dominate…