1 paper
Jaerin Lee, Kyoung Mu Lee
Recent works have shown that gradient-update alignment is a powerful signal for modulating optimizer updates, often leading to faster training. We promote this update-wise heuristi…