1 paper
Kang Liu, Suyan Li
A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimizer states. We study this delayed effect th…