activity
20182022
most citedRevisiting Model Stitching to Compare Neural Representations

22 citations · 30 across the 2 of their papers we have counts for

collaborators

6 papers

cs.LG20228 cited

Data Scaling Laws in NMT: The Effect of Noise and Architecture

Yamini Bansal, Behrooz Ghorbani, Ankush Garg +5

In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that…

cs.LG202122 cited

Revisiting Model Stitching to Compare Neural Representations

Yamini Bansal, Preetum Nakkiran, Boaz Barak

We revisit and extend model stitching (Lenc & Vedaldi 2015) as a methodology to study the internal representations of neural networks. Given two trained and frozen models and $…

cs.LG2020

For self-supervised learning, Rationality implies generalization, provably

Yamini Bansal, Gal Kaplun, Boaz Barak

We prove a new upper bound on the generalization gap of classifiers that are obtained by first using self-supervision to learn a representation of the training data, and then f…

cs.LG2020

Distributional Generalization: A New Kind of Generalization

Preetum Nakkiran, Yamini Bansal

We introduce a new notion of generalization -- Distributional Generalization -- which roughly states that outputs of a classifier at train and test time are close *as distributions…

cs.LG2019

Deep Double Descent: Where Bigger Models and More Data Hurt

Preetum Nakkiran, Gal Kaplun, Yamini Bansal +3

We show that a variety of modern deep learning tasks exhibit a "double-descent" phenomenon where, as we increase model size, performance first gets worse and then gets better. More…

stat.ML2018

Minnorm training: an algorithm for training over-parameterized deep neural networks

Yamini Bansal, Madhu Advani, David D Cox +1

In this work, we propose a new training method for finding minimum weight norm solutions in over-parameterized neural networks (NNs). This method seeks to improve training speed an…