2 papers
cs.LG2023
Layer-wise Adaptive Step-Sizes for Stochastic First-Order Methods for Deep Learning
Achraf Bahamou, Donald Goldfarb
We propose a new per-layer adaptive step-size procedure for stochastic first-order optimization methods for minimizing empirical loss functions in deep learning, eliminating the ne…
cs.LG2022
A Mini-Block Fisher Method for Deep Neural Networks
Achraf Bahamou, Donald Goldfarb, Yi Ren
Deep neural networks (DNNs) are currently predominantly trained using first-order methods. Some of these methods (e.g., Adam, AdaGrad, and RMSprop, and their variants) incorporate…