activity
20172022
most citedFine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks

256 citations · 439 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

19 papers · 1 filter

cs.LG2023

Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature Connectivity

Zhanpeng Zhou, Yongyi Yang, Xiaojiang Yang +2

Recent work has revealed many intriguing empirical phenomena in neural network training, despite the poorly understood and highly complex loss landscapes and training dynamics. One…

cs.LG2023

Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations

Yongyi Yang, Jacob Steinhardt, Wei Hu

Recent work has observed an intriguing ''Neural Collapse'' phenomenon in well-trained neural networks, where the last-layer representations of training samples with the same label…

cs.LG20225 cited

Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data

Spencer Frei, Gal Vardi, Peter L. Bartlett +2

The implicit biases of gradient-based optimization algorithms are conjectured to be a major factor in the success of modern deep learning. In this work, we investigate the implicit…

cs.LG20228 cited

More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize

Alexander Wei, Wei Hu, Jacob Steinhardt

Of theories for why large-scale machine learning models generalize despite being vastly overparameterized, which of their assumptions are needed to capture the qualitative phenomen…

cs.LG202013 cited

The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks

Wei Hu, Lechao Xiao, Ben Adlam +1

Modern neural networks are often regarded as complex black-box functions whose behavior is difficult to understand owing to their nonlinear dependence on the data and the nonconvex…

cs.LG2020

Few-Shot Learning via Learning the Representation, Provably

Simon S. Du, Wei Hu, Sham M. Kakade +2

This paper studies few-shot learning via representation learning, where one uses source tasks with data per task to learn a representation in order to reduce the sample c…