2 citations · 3 across the 4 of their papers we have counts for
1 paper · 1 filter
Tianle Cai, Ruiqi Gao, Jikai Hou +5
First-order methods such as stochastic gradient descent (SGD) are currently the standard algorithm for training deep neural networks. Second-order methods, despite their better con…