activity
20102019
most citedOn the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization

146 citations · 286 across the 7 of their papers we have counts for

collaborators
Showing math.OCShow all

9 papers · 1 filter

math.OC2019

Online Primal-Dual Mirror Descent under Stochastic Constraints

Xiaohan Wei, Hao Yu, Michael J. Neely

We consider online convex optimization with stochastic constraints where the objective functions are arbitrarily time-varying and the constraint functions are independent and ident…

math.OC201931 cited

On the Computation and Communication Complexity of Parallel SGD with Dynamic Batch Sizes for Stochastic Non-Convex Optimization

Hao Yu, Rong Jin

For SGD based distributed stochastic optimization, computation complexity, measured by the convergence rate in terms of the number of stochastic gradient calls, and communication c…

math.OC2019146 cited

On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization

Hao Yu, Rong Jin, Sen Yang

Recent developments on large-scale distributed machine learning applications, e.g., deep neural networks, benefit enormously from the advances in distributed non-convex optimizatio…

math.OC2018

Solving Non-smooth Constrained Programs with Lower Complexity than : A Primal-Dual Homotopy Smoothing Approach

Xiaohan Wei, Hao Yu, Qing Ling +1

We propose a new primal-dual homotopy smoothing algorithm for a linearly constrained convex program, where neither the primal nor the dual function has to be smooth or strongly con…

math.OC2018

Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning

Hao Yu, Sen Yang, Shenghuo Zhu

In distributed training of deep neural networks, parallel mini-batch SGD is widely used to speed up the training process by using multiple workers. It uses multiple workers to samp…

math.OC20171 cited

Online Learning in Weakly Coupled Markov Decision Processes: A Convergence Time Study

Xiaohan Wei, Hao Yu, Michael J. Neely

We consider multiple parallel Markov decision processes (MDPs) coupled by global constraints, where the time varying objective and constraint functions can only be observed after t…