146 citations · 192 across the 9 of their papers we have counts for
4 papers · 1 filter
On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization
Hao Yu, Rong Jin, Sen Yang
Recent developments on large-scale distributed machine learning applications, e.g., deep neural networks, benefit enormously from the advances in distributed non-convex optimizatio…
Shrinking the Upper Confidence Bound: A Dynamic Product Selection Problem for Urban Warehouses
Rong Jin, David Simchi-Levi, Li Wang +2
The recent rising popularity of ultra-fast delivery services on retail platforms fuels the increasing use of urban warehouses, whose proximity to customers makes fast deliveries vi…
On the Convergence of (Stochastic) Gradient Descent with Extrapolation for Non-Convex Optimization
Yi Xu, Zhuoning Yuan, Sen Yang +2
Extrapolation is a well-known technique for solving convex optimization and variational inequalities and recently attracts some attention for non-convex optimization. Several recen…
Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning
Hao Yu, Sen Yang, Shenghuo Zhu
In distributed training of deep neural networks, parallel mini-batch SGD is widely used to speed up the training process by using multiple workers. It uses multiple workers to samp…