activity
20152022
most citedUn-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization

67 citations · 110 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2022★ 9 cited

Learning from many trajectories

Stephen Tu, Roy Frostig, Mahdi Soltanolkotabi

We initiate a study of supervised learning from many independent sequences ("trajectories") of non-independent covariates, reflecting tasks in sequence modeling, control, and reinf…

cs.LG2021

Efficient and Modular Implicit Differentiation

Mathieu Blondel, Quentin Berthet, Marco Cuturi +5

Automatic differentiation (autodiff) has revolutionized machine learning. It allows to express complex computations by composing elementary ones in creative ways and removes the bu…

cs.LG2019★ 14 cited

The advantages of multiple classes for reducing overfitting from test set reuse

Vitaly Feldman, Roy Frostig, Moritz Hardt

Excessive reuse of holdout data can lead to overfitting. However, there is little concrete evidence of significant overfitting due to holdout reuse in popular multiclass benchmarks…

cs.LG2018

Measuring the Effects of Data Parallelism on Neural Network Training

Christopher J. Shallue, Jaehoon Lee, Joseph Antognini +3

Recent hardware developments have dramatically increased the scale of data parallelism available for neural network training. Among the simplest ways to harness next-generation har…

cs.LG2017★ 7 cited

Random Features for Compositional Kernels

Amit Daniely, Roy Frostig, Vineet Gupta +1

We describe and analyze a simple random feature scheme (RFS) from prescribed compositional kernels. The compositional kernels we use are inspired by the structure of convolutional…