activity
20172023
most citedA Unified Approach to Adaptive Regularization in Online and Stochastic Optimization

9 citations · 17 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG20211 cited

Large-Scale Differentially Private BERT

Rohan Anil, Badih Ghazi, Vineet Gupta +2

In this work, we study the large-scale pretraining of BERT-Large with differentially private SGD (DP-SGD). We show that combined with a careful implementation, scaling up the batch…

cs.LG2020

Scalable Second Order Optimization for Deep Learning

Rohan Anil, Vineet Gupta, Tomer Koren +2

Optimization in machine learning, both theoretical and applied, is presently dominated by first-order gradient methods such as stochastic gradient descent. Second-order optimizatio…

cs.LG2019

Memory-Efficient Adaptive Optimization

Rohan Anil, Vineet Gupta, Tomer Koren +1

Adaptive gradient-based optimizers such as Adagrad and Adam are crucial for achieving state-of-the-art performance in machine translation and language modeling. However, these meth…

cs.LG2018

The Singular Values of Convolutional Layers

Hanie Sedghi, Vineet Gupta, Philip M. Long

We characterize the singular values of the linear transformation associated with a standard 2D multi-channel convolutional layer, enabling their efficient computation. This charact…

cs.LG2018

Shampoo: Preconditioned Stochastic Tensor Optimization

Vineet Gupta, Tomer Koren, Yoram Singer

Preconditioned gradient methods are among the most general and powerful tools in optimization. However, preconditioning requires storing and manipulating prohibitively large matric…

cs.LG20179 cited

A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization

Vineet Gupta, Tomer Koren, Yoram Singer

We describe a framework for deriving and analyzing online optimization algorithms that incorporate adaptive, data-dependent regularization, also termed preconditioning. Such algori…