2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2024
Narrowing the Focus: Learned Optimizers for Pretrained Models
Gus Kristiansen, Mark Sandler, Andrey Zhmoginov +4
In modern deep learning, the models are learned by applying gradient updates using an optimizer, which transforms the updates based on various statistics. Optimizers are often hand…
cs.LG2023★ 2 cited
Training trajectories, mini-batch losses and the curious role of the learning rate
Mark Sandler, Andrey Zhmoginov, Max Vladymyrov +1
Stochastic gradient descent plays a fundamental role in nearly all applications of deep learning. However its ability to converge to a global minimum remains shrouded in mystery. I…