activity
20122022
most citedAugment your batch: better training with larger batches

50 citations · 171 across the 12 of their papers we have counts for

collaborators
Showing 2019Show all

10 papers · 1 filter

cs.LG2019

Is Feature Diversity Necessary in Neural Network Initialization?

Yaniv Blumenfeld, Dar Gilboa, Daniel Soudry

Standard practice in training neural networks involves initializing the weights in an independent fashion. The results of recent work suggest that feature "diversity" at initializa…

cs.LG2019

The Knowledge Within: Methods for Data-Free Model Compression

Matan Haroush, Itay Hubara, Elad Hoffer +1

Recently, an extensive amount of research has been focused on compressing and accelerating Deep Neural Networks (DNN). So far, high compression rate algorithms require part of the…

cs.LG201916 cited

A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case

Greg Ongie, Rebecca Willett, Daniel Soudry +1

A key element of understanding the efficacy of overparameterized neural networks is characterizing how they represent functions as the number of weights in the network approaches i…

cs.LG2019

At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?

Niv Giladi, Mor Shpigel Nacson, Elad Hoffer +1

Background: Recent developments have made it possible to accelerate neural networks training significantly using large batch sizes and data parallelism. Training in an asynchronous…

cs.CV2019

Mix & Match: training convnets with mixed image sizes for improved accuracy, speed and scale resiliency

Elad Hoffer, Berry Weinstein, Itay Hubara +3

Convolutional neural networks (CNNs) are commonly trained using a fixed spatial image size predetermined for a given model. Although trained on images of aspecific size, it is well…

cs.LG2019

Kernel and Rich Regimes in Overparametrized Models

Blake Woodworth, Suriya Gunasekar, Pedro Savarese +5

A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus trai…