1 paper · 1 filter
Oscar Smee, Fred Roosta, Stephen J. Wright
Gradient descent is the primary workhorse for optimizing large-scale problems in machine learning. However, its performance is highly sensitive to the choice of the learning rate.…