1 paper
Xianliang Xu, Ting Du, Wang Kong +3
In the context of over-parameterization, there is a line of work demonstrating that randomly initialized (stochastic) gradient descent (GD) converges to a globally optimal solution…