117 citations · 147 across the 8 of their papers we have counts for
1 paper · 2 filters
Yiping Lu, Chao Ma, Yulong Lu +2
Training deep neural networks with stochastic gradient descent (SGD) can often achieve zero training loss on real-world tasks although the optimization landscape is known to be hig…