activity
20162021
most citedWide & Deep Learning for Recommender Systems

263 citations · 485 across the 5 of their papers we have counts for

collaborators

15 papers

cs.LG20227 cited

Learning from Randomly Initialized Neural Network Features

Ehsan Amid, Rohan Anil, Wojciech Kotłowski +1

We present the surprising result that randomly initialized neural networks are good feature extractors in expectation. These random features correspond to finite-sample realization…

cs.LG20224 cited

Step-size Adaptation Using Exponentiated Gradient Updates

Ehsan Amid, Rohan Anil, Christopher Fifty +1

Optimizers like Adam and AdaGrad have been very successful in training large-scale neural networks. Yet, the performance of these methods is heavily dependent on a carefully tuned…

cs.LG202129 cited

Efficiently Identifying Task Groupings for Multi-Task Learning

Christopher Fifty, Ehsan Amid, Zhe Zhao +3

Multi-task learning can leverage information learned by one task to benefit the training of other tasks. Despite this capacity, naively training all tasks together in one model oft…

cs.LG20211 cited

Large-Scale Differentially Private BERT

Rohan Anil, Badih Ghazi, Vineet Gupta +2

In this work, we study the large-scale pretraining of BERT-Large with differentially private SGD (DP-SGD). We show that combined with a careful implementation, scaling up the batch…

cs.LG2021

A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes

Zachary Nado, Justin M. Gilmer, Christopher J. Shallue +2

Recently the LARS and LAMB optimizers have been proposed for training neural networks faster using large batch sizes. LARS and LAMB add layer-wise normalization to the update rules…

cs.LG2020

Stochastic Optimization with Laggard Data Pipelines

Naman Agarwal, Rohan Anil, Tomer Koren +2

State-of-the-art optimization is steadily shifting towards massively parallel pipelines with extremely large batch sizes. As a consequence, CPU-bound preprocessing and disk/memory/…