1 paper
Guoxiang Xu, Bince Qu, Qi Sun +1
Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight, jointly changing its norm and…