8 citations · 8 across the 1 of their papers we have counts for
2 papers
math.OC2020
Bounding the expected run-time of nonconvex optimization with early stopping
Thomas Flynn, Kwang Min Yu, Abid Malik +2
This work examines the convergence of stochastic gradient-based optimization algorithms that use early stopping based on a validation function. The form of early stopping we consid…
cs.LG2019★ 8 cited
Layered SGD: A Decentralized and Synchronous SGD Algorithm for Scalable Deep Neural Network Training
Kwangmin Yu, Thomas Flynn, Shinjae Yoo +1
Stochastic Gradient Descent (SGD) is the most popular algorithm for training deep neural networks (DNNs). As larger networks and datasets cause longer training times, training on d…