16 citations · 16 across the 1 of their papers we have counts for
1 paper · 1 filter
Tianle Cai, Ruiqi Gao, Jikai Hou +5
First-order methods such as stochastic gradient descent (SGD) are currently the standard algorithm for training deep neural networks. Second-order methods, despite their better con…