112 citations · 288 across the 11 of their papers we have counts for
12 papers · 1 filter
Solving Non-Convex Non-Concave Min-Max Games Under Polyak-Łojasiewicz Condition
Maziar Sanjabi, Meisam Razaviyayn, Jason D. Lee
In this short note, we consider the problem of solving a min-max zero-sum game. This problem has been extensively studied in the convex-concave regime where the global solution can…
Gradient Descent Finds Global Minima of Deep Neural Networks
Simon S. Du, Jason D. Lee, Haochuan Li +2
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero tr…
Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel
Colin Wei, Jason D. Lee, Qiang Liu +1
Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the…
Convergence to Second-Order Stationarity for Constrained Non-Convex Optimization
Maher Nouiehed, Jason D. Lee, Meisam Razaviyayn
We consider the problem of finding an approximate second-order stationary point of a constrained non-convex optimization problem. We first show that, unlike the gradient descent me…
Provably Correct Automatic Subdifferentiation for Qualified Programs
Sham Kakade, Jason D. Lee
The Cheap Gradient Principle (Griewank 2008) --- the computational cost of computing the gradient of a scalar-valued function is nearly the same (often within a factor of ) as t…
Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
Simon S. Du, Wei Hu, Jason D. Lee
We study the implicit regularization imposed by gradient descent for learning multi-layer homogeneous functions including feed-forward fully connected and convolutional deep neural…