activity
20172022
most citedShape Matters: Understanding the Implicit Bias of the Noise Covariance

18 citations · 21 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20222 cited

Beyond Separability: Analyzing the Linear Transferability of Contrastive Representations to Related Subpopulations

Jeff Z. HaoChen, Colin Wei, Ananya Kumar +1

Contrastive learning is a highly effective method for learning representations from unlabeled data. Recent works show that contrastive representations can transfer across domains,…

cs.LG20201 cited

Meta-learning Transferable Representations with a Single Target Domain

Hong Liu, Jeff Z. HaoChen, Colin Wei +1

Recent works found that fine-tuning and joint training---two popular approaches for transfer learning---do not always improve accuracy on downstream tasks. First, we aim to underst…

cs.LG202018 cited

Shape Matters: Understanding the Implicit Bias of the Noise Covariance

Jeff Z. HaoChen, Colin Wei, Jason D. Lee +1

The noise in stochastic gradient descent (SGD) provides a crucial implicit regularization effect for training overparameterized models. Prior theoretical work largely focuses on sp…

cs.LG2020

Self-training Avoids Using Spurious Features Under Domain Shift

Yining Chen, Colin Wei, Ananya Kumar +1

In unsupervised domain adaptation, existing theory focuses on situations where the source and target domains are close. In practice, conditional entropy minimization and pseudo-lab…

cs.LG2020

The Implicit and Explicit Regularization Effects of Dropout

Colin Wei, Sham Kakade, Tengyu Ma

Dropout is a widely-used regularization technique, often required to obtain state-of-the-art for a number of architectures. This work demonstrates that dropout introduces two disti…

cs.LG2019

Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks

Yuanzhi Li, Colin Wei, Tengyu Ma

Stochastic gradient descent with a large initial learning rate is widely used for training modern neural net architectures. Although a small initial learning rate allows for faster…