1 paper · 1 filter
Mitchell Scott, Tianshi Xu, Ziyuan Tang +4
Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. We analyze preconditioned SGD in the geometry induced b…