27 citations · 62 across the 8 of their papers we have counts for
15 papers
A Variance-Reduced Stochastic Accelerated Primal Dual Algorithm
Bugra Can, Mert Gurbuzbalaban, Necdet Serhat Aybat
In this work, we consider strongly convex strongly concave (SCSC) saddle point (SP) problems where is -smooth,…
L-DQN: An Asynchronous Limited-Memory Distributed Quasi-Newton Method
Bugra Can, Saeed Soori, Maryam Mehri Dehnavi +1
This work proposes a distributed algorithm for solving empirical risk minimization problems, called L-DQN, under the master/worker communication model. L-DQN is a distributed limit…
Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
Alexander Camuto, Xiaoyu Wang, Lingjiong Zhu +3
Gaussian noise injections (GNIs) are a family of simple and widely-used regularisation methods for training neural networks, where one injects additive or multiplicative Gaussian n…
IDEAL: Inexact DEcentralized Accelerated Augmented Lagrangian Method
Yossi Arjevani, Joan Bruna, Bugra Can +3
We introduce a framework for designing primal methods under the decentralized optimization setting where local functions are smooth and strongly convex. Our approach consists of ap…
Fractional moment-preserving initialization schemes for training deep neural networks
Mert Gurbuzbalaban, Yuanhan Hu
A traditional approach to initialization in deep neural networks (DNNs) is to sample the network weights randomly for preserving the variance of pre-activations. On the other hand,…
Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise
Umut Şimşekli, Lingjiong Zhu, Yee Whye Teh +1
Stochastic gradient descent with momentum (SGDm) is one of the most popular optimization algorithms in deep learning. While there is a rich theory of SGDm for convex problems, the…