4 papers
Training Neural Networks at Any Scale
Thomas Pethick, Kimon Antonakopoulos, Antonio Silveti-Falls +2
This article reviews modern optimization methods for training neural networks with an emphasis on efficiency and scale. We present state-of-the-art optimization algorithms under a…
Adaptive Conditional Gradient Descent
Abbas Khademi, Antonio Silveti-Falls
Selecting an effective step-size is a fundamental challenge in first-order optimization, especially for problems with non-Euclidean geometries. This paper presents a novel adaptive…
Generalized Gradient Norm Clipping & Non-Euclidean -Smoothness
Thomas Pethick, Wanyun Xie, Mete Erdogan +3
This work introduces a hybrid non-Euclidean optimization method which generalizes gradient norm clipping by combining steepest descent and conditional gradient approaches. The meth…
Training Deep Learning Models with Norm-Constrained LMOs
Thomas Pethick, Wanyun Xie, Kimon Antonakopoulos +3
In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO…