48 citations · 63 across the 5 of their papers we have counts for
7 papers
Knowledge Distillation: Bad Models Can Be Good Role Models
Gal Kaplun, Eran Malach, Preetum Nakkiran +1
Large neural networks trained in the overparameterized regime are able to fit noise to zero train error. Recent work \citep{nakkiran2020distributional} has empirically observed tha…
For Manifold Learning, Deep Neural Networks can be Locality Sensitive Hash Functions
Nishanth Dikkala, Gal Kaplun, Rina Panigrahy
It is well established that training deep neural networks gives useful representations that capture essential features of the inputs. However, these representations are poorly unde…
For self-supervised learning, Rationality implies generalization, provably
Yamini Bansal, Gal Kaplun, Boaz Barak
We prove a new upper bound on the generalization gap of classifiers that are obtained by first using self-supervision to learn a representation of the training data, and then f…
Robustness from Simple Classifiers
Sharon Qian, Dimitris Kalimeris, Gal Kaplun +1
Despite the vast success of Deep Neural Networks in numerous application domains, it has been shown that such models are not robust i.e., they are vulnerable to small adversarial p…
Deep Double Descent: Where Bigger Models and More Data Hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal +3
We show that a variety of modern deep learning tasks exhibit a "double-descent" phenomenon where, as we increase model size, performance first gets worse and then gets better. More…
SGD on Neural Networks Learns Functions of Increasing Complexity
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris +4
We perform an experimental study of the dynamics of Stochastic Gradient Descent (SGD) in learning deep neural networks for several real and synthetic classification tasks. We show…