2 papers
cs.LG2025
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
Aaron Mishkin, Ahmed Khaled, Yuanhao Wang +2
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case con…
cs.LG2024
The Road Less Scheduled
Aaron Defazio, Xingyu Alice Yang, Harsh Mehta +3
Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We pro…