1 paper
Runzhe Wang, Sadhika Malladi, Tianhao Wang +2
Momentum is known to accelerate the convergence of gradient descent in strongly convex settings without stochastic gradient noise. In stochastic optimization, such as training neur…