23 citations · 26 across the 3 of their papers we have counts for
4 papers
On the Generalization Mystery in Deep Learning
Satrajit Chatterjee, Piotr Zielinski
The generalization mystery in deep learning is the following: Why do over-parameterized neural networks trained with gradient descent (GD) generalize well on real datasets even tho…
Enabling Binary Neural Network Training on the Edge
Erwei Wang, James J. Davis, Daniele Moro +6
The ever-growing computational demands of increasingly complex machine learning models frequently necessitate the use of powerful cloud-based infrastructure for their training. Bin…
Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
Satrajit Chatterjee, Piotr Zielinski
We propose a new metric (-coherence) to experimentally study the alignment of per-example gradients during training. Intuitively, given a sample of size , -coherence is th…
Weak and Strong Gradient Directions: Explaining Memorization, Generalization, and Hardness of Examples at Scale
Piotr Zielinski, Shankar Krishnan, Satrajit Chatterjee
Coherent Gradients (CGH) is a recently proposed hypothesis to explain why over-parameterized neural networks trained with gradient descent generalize well even though they have suf…