2 citations · 3 across the 5 of their papers we have counts for
1 paper · 1 filter
Kaelan Donatella, Samuel Duffield, Maxwell Aifer +3
Second-order training methods have better convergence properties than gradient descent but are rarely used in practice for large-scale training due to their computational overhead.…