229 citations · 547 across the 23 of their papers we have counts for
9 papers · 1 filter
Non-Gaussianity of Stochastic Gradient Noise
Abhishek Panigrahi, Raghav Somani, Navin Goyal +1
What enables Stochastic Gradient Descent (SGD) to achieve better generalization than Gradient Descent (GD) in Neural Network training? This question has attracted much attention. I…
Efficient Algorithms for Smooth Minimax Optimization
Kiran Koshy Thekumparampil, Prateek Jain, Praneeth Netrapalli +1
This paper studies first order methods for solving smooth minimax optimization problems where is smooth and is concave for each…
Making the Last Iterate of SGD Information Theoretically Optimal
Prateek Jain, Dheeraj Nagaraj, Praneeth Netrapalli
Stochastic gradient descent (SGD) is one of the most widely used algorithms for large scale optimization problems. While classical theoretical analysis of SGD for convex problems s…
The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure For Least Squares
Rong Ge, Sham M. Kakade, Rahul Kidambi +1
Minimax optimal convergence rates for classes of stochastic convex optimization problems are well characterized, where the majority of results utilize iterate averaged stochastic g…
Online Non-Convex Learning: Following the Perturbed Leader is Optimal
Arun Sai Suggala, Praneeth Netrapalli
We study the problem of online learning with non-convex losses, where the learner has access to an offline optimization oracle. We show that the classical Follow the Perturbed Lead…
SGD without Replacement: Sharper Rates for General Smooth Convex Functions
Prateek Jain, Dheeraj Nagaraj, Praneeth Netrapalli
We study stochastic gradient descent {\em without replacement} (\sgdwor) for smooth convex functions. \sgdwor is widely observed to converge faster than true \sgd where each sample…