activity
20162023
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 696 across the 19 of their papers we have counts for

collaborators
Showing 2020Show all

7 papers · 1 filter

cs.LG2020

Parallel Training of Deep Networks with Local Updates

Michael Laskin, Luke Metz, Seth Nabarro +5

Deep learning models trained on large data sets have been widely successful in both vision and language domains. As state-of-the-art deep learning architectures have continued to g…

cs.LG2020

Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian

Jack Parker-Holder, Luke Metz, Cinjon Resnick +6

Over the last decade, a single algorithm has changed many facets of our lives - Stochastic Gradient Descent (SGD). In the era of ever decreasing loss functions, SGD and its various…

cs.LG2020

Reverse engineering learned optimizers reveals known and novel mechanisms

Niru Maheswaranathan, David Sussillo, Luke Metz +2

Learned optimizers are algorithms that can themselves be trained to solve optimization problems. In contrast to baseline optimizers (such as momentum or Adam) that use simple updat…

cs.LG2020★ 14 cited

Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves

Luke Metz, Niru Maheswaranathan, C. Daniel Freeman +2

Much as replacing hand-designed features with learned functions has revolutionized how we solve perceptual tasks, we believe learned algorithms will transform how we train models.…

stat.ML2020★ 16 cited

On Linear Identifiability of Learned Representations

Geoffrey Roeder, Luke Metz, Diederik P. Kingma

Identifiability is a desirable property of a statistical model: it implies that the true model parameters may be estimated to any desired precision, given sufficient computational…

cs.LG2020

Using a thousand optimization tasks to learn hyperparameter search strategies

Luke Metz, Niru Maheswaranathan, Ruoxi Sun +3

We present TaskSet, a dataset of tasks for use in training and evaluating optimizers. TaskSet is unique in its size and diversity, containing over a thousand tasks ranging from ima…