1 paper
Junho So, Dongwook Shin
During SGD training, the gradients often align strongly with the dominant subspace spanned by the top-k eigenvectors of the Hessian of the loss. While this seems to naturally imp…