256 citations · 439 across the 8 of their papers we have counts for
19 papers · 1 filter
Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature Connectivity
Zhanpeng Zhou, Yongyi Yang, Xiaojiang Yang +2
Recent work has revealed many intriguing empirical phenomena in neural network training, despite the poorly understood and highly complex loss landscapes and training dynamics. One…
Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations
Yongyi Yang, Jacob Steinhardt, Wei Hu
Recent work has observed an intriguing ''Neural Collapse'' phenomenon in well-trained neural networks, where the last-layer representations of training samples with the same label…
Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data
Spencer Frei, Gal Vardi, Peter L. Bartlett +2
The implicit biases of gradient-based optimization algorithms are conjectured to be a major factor in the success of modern deep learning. In this work, we investigate the implicit…
More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize
Alexander Wei, Wei Hu, Jacob Steinhardt
Of theories for why large-scale machine learning models generalize despite being vastly overparameterized, which of their assumptions are needed to capture the qualitative phenomen…
The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
Wei Hu, Lechao Xiao, Ben Adlam +1
Modern neural networks are often regarded as complex black-box functions whose behavior is difficult to understand owing to their nonlinear dependence on the data and the nonconvex…
Few-Shot Learning via Learning the Representation, Provably
Simon S. Du, Wei Hu, Sham M. Kakade +2
This paper studies few-shot learning via representation learning, where one uses source tasks with data per task to learn a representation in order to reduce the sample c…