306 citations · 338 across the 8 of their papers we have counts for
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2018
Differential Equations for Modeling Asynchronous Algorithms
Li He, Qi Meng, Wei Chen +2
Asynchronous stochastic gradient descent (ASGD) is a popular parallel optimization algorithm in machine learning. Most theoretical analysis on ASGD take a discrete view and prove u…
stat.ML2018
-SGD: Optimizing ReLU Neural Networks in its Positively Scale-Invariant Space
Qi Meng, Shuxin Zheng, Huishuai Zhang +3
It is well known that neural networks with rectified linear units (ReLU) activation functions are positively scale-invariant. Conventional algorithms like stochastic gradient desce…
stat.ML2017★ 18 cited
Convergence Analysis of Distributed Stochastic Gradient Descent with Shuffling
Qi Meng, Wei Chen, Yue Wang +2
When using stochastic gradient descent to solve large-scale machine learning problems, a common practice of data processing is to shuffle the training data, partition the data acro…