9 citations · 17 across the 4 of their papers we have counts for
7 papers · 1 filter
Large-Scale Differentially Private BERT
Rohan Anil, Badih Ghazi, Vineet Gupta +2
In this work, we study the large-scale pretraining of BERT-Large with differentially private SGD (DP-SGD). We show that combined with a careful implementation, scaling up the batch…
Scalable Second Order Optimization for Deep Learning
Rohan Anil, Vineet Gupta, Tomer Koren +2
Optimization in machine learning, both theoretical and applied, is presently dominated by first-order gradient methods such as stochastic gradient descent. Second-order optimizatio…
Memory-Efficient Adaptive Optimization
Rohan Anil, Vineet Gupta, Tomer Koren +1
Adaptive gradient-based optimizers such as Adagrad and Adam are crucial for achieving state-of-the-art performance in machine translation and language modeling. However, these meth…
The Singular Values of Convolutional Layers
Hanie Sedghi, Vineet Gupta, Philip M. Long
We characterize the singular values of the linear transformation associated with a standard 2D multi-channel convolutional layer, enabling their efficient computation. This charact…
Shampoo: Preconditioned Stochastic Tensor Optimization
Vineet Gupta, Tomer Koren, Yoram Singer
Preconditioned gradient methods are among the most general and powerful tools in optimization. However, preconditioning requires storing and manipulating prohibitively large matric…
A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization
Vineet Gupta, Tomer Koren, Yoram Singer
We describe a framework for deriving and analyzing online optimization algorithms that incorporate adaptive, data-dependent regularization, also termed preconditioning. Such algori…