20 citations · 22 across the 2 of their papers we have counts for
4 papers
Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
Satrajit Chatterjee, Piotr Zielinski
We propose a new metric (-coherence) to experimentally study the alignment of per-example gradients during training. Intuitively, given a sample of size , -coherence is th…
Weak and Strong Gradient Directions: Explaining Memorization, Generalization, and Hardness of Examples at Scale
Piotr Zielinski, Shankar Krishnan, Satrajit Chatterjee
Coherent Gradients (CGH) is a recently proposed hypothesis to explain why over-parameterized neural networks trained with gradient descent generalize well even though they have suf…
Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
Satrajit Chatterjee
An open question in the Deep Learning community is why neural networks trained with Gradient Descent generalize well on real datasets even though they are capable of fitting random…
Circuit-Based Intrinsic Methods to Detect Overfitting
Satrajit Chatterjee, Alan Mishchenko
The focus of this paper is on intrinsic methods to detect overfitting. By intrinsic methods, we mean methods that rely only on the model and the training data, as opposed to tradit…