activity
20122022
most citedAugment your batch: better training with larger batches

50 citations · 171 across the 12 of their papers we have counts for

collaborators
Showing 2018Show all

9 papers · 1 filter

cs.CV2018

Post-training 4-bit quantization of convolution networks for rapid-deployment

Ron Banner, Yury Nahshan, Elad Hoffer +1

Convolutional neural networks require significant memory bandwidth and storage for intermediate computations, apart from substantial computing resources. Neural network quantizatio…

cs.LG2018

Scalable Methods for 8-bit Training of Neural Networks

Ron Banner, Itay Hubara, Elad Hoffer +1

Quantized Neural Networks (QNNs) are often used to improve network efficiency during the inference phase, i.e. after the network has been trained. Extensive research in the field s…

cs.LG2018

Implicit Bias of Gradient Descent on Linear Convolutional Networks

Suriya Gunasekar, Jason Lee, Daniel Soudry +1

We show that gradient descent on full-width linear convolutional networks of depth converges to a linear predictor related to the bridge penalty in the frequency d…

cs.LG2018

The Global Optimization Geometry of Shallow Linear Neural Networks

Zhihui Zhu, Daniel Soudry, Yonina C. Eldar +1

We examine the squared error loss landscape of shallow linear neural networks. We show---with significantly milder assumptions than previous works---that the corresponding optimiza…

stat.ML2018

Task Agnostic Continual Learning Using Online Variational Bayes

Chen Zeno, Itay Golan, Elad Hoffer +1

Catastrophic forgetting is the notorious vulnerability of neural networks to the change of the data distribution while learning. This phenomenon has long been considered a major ob…

stat.ML2018

Convergence of Gradient Descent on Separable Data

Mor Shpigel Nacson, Jason D. Lee, Suriya Gunasekar +3

We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets…