168 citations · 172 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023★ 4 cited
Why Does Sharpness-Aware Minimization Generalize Better Than SGD?
Zixiang Chen, Junkai Zhang, Yiwen Kou +3
The challenge of overfitting, in which the model memorizes the training data and fails to generalize to test data, has become increasingly significant in the training of large neur…
cs.LG2023★ 168 cited
Symbolic Discovery of Optimization Algorithms
Xiangning Chen, Chen Liang, Da Huang +9
We present a method to formulate algorithm discovery as program search, and apply it to discover optimization algorithms for deep neural network training. We leverage efficient sea…