20 citations · 23 across the 2 of their papers we have counts for
2 papers
cs.LG2020★ 20 cited
AdaScale SGD: A User-Friendly Algorithm for Distributed Training
Tyler B. Johnson, Pulkit Agrawal, Haijie Gu +1
When using large-batch training to speed up stochastic gradient descent, learning rates must adapt to new batch sizes in order to maximize speed-ups and preserve model quality. Re-…
stat.ME2012★ 3 cited
Sequential Nonparametric Regression
Haijie Gu, John Lafferty
We present algorithms for nonparametric regression in settings where the data are obtained sequentially. While traditional estimators select bandwidths that depend upon the sample…