8 papers
Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex Problems
Stefano Sarao Mannelli, Pierfrancesco Urbani
The optimization step in many machine learning problems rarely relies on vanilla gradient descent but it is common practice to use momentum-based accelerated methods. Despite these…
Post-Workshop Report on Science meets Engineering in Deep Learning, NeurIPS 2019, Vancouver
Levent Sagun, Caglar Gulcehre, Adriana Romero +2
Science meets Engineering in Deep Learning took place in Vancouver as part of the Workshop section of NeurIPS 2019. As organizers of the workshop, we created the following report i…
Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota +3
Despite the widespread use of gradient-based algorithms for optimizing high-dimensional non-convex functions, understanding their ability of finding good minima instead of being tr…
Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka Zdeborová
We study the dynamics of optimization and the generalization properties of one-hidden layer neural networks with quadratic activation function in the over-parametrized regime where…
Thresholds of descending algorithms in inference problems
Stefano Sarao Mannelli, Lenka Zdeborova
We review recent works on analyzing the dynamics of gradient-based algorithms in a prototypical statistical inference problem. Using methods and insights from the physics of glassy…
Who is Afraid of Big Bad Minima? Analysis of Gradient-Flow in a Spiked Matrix-Tensor Model
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota +2
Gradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in op…