1 paper
Matteo Pagliardini, Pierre Ablin, David Grangier
Momentum based optimizers are central to a wide range of machine learning applications. These typically rely on an Exponential Moving Average (EMA) of gradients, which decays expon…