23 citations · 31 across the 5 of their papers we have counts for
4 papers
Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal Data
Yiwen Kou, Zixiang Chen, Quanquan Gu
The implicit bias towards solutions with favorable properties is believed to be a key reason why neural networks trained by gradient-based optimization can generalize well. While t…
Why Does Sharpness-Aware Minimization Generalize Better Than SGD?
Zixiang Chen, Junkai Zhang, Yiwen Kou +3
The challenge of overfitting, in which the model memorizes the training data and fails to generalize to test data, has become increasingly significant in the training of large neur…
Probing Heavy Neutrinos at the LHC from Fat-jet using Machine Learning
Wei Liu, Jing Li, Zixiang Chen +1
We explore the potential to use machine learning methods to search for heavy neutrinos, from their hadronic final states including a fat-jet signal, via the processes $pp \rightarr…
Towards Understanding Mixture of Experts in Deep Learning
Zixiang Chen, Yihe Deng, Yue Wu +2
The Mixture-of-Experts (MoE) layer, a sparsely-activated model controlled by a router, has achieved great success in deep learning. However, the understanding of such architecture…