33 citations · 55 across the 7 of their papers we have counts for
4 papers · 1 filter
Why Does Sharpness-Aware Minimization Generalize Better Than SGD?
Zixiang Chen, Junkai Zhang, Yiwen Kou +3
The challenge of overfitting, in which the model memorizes the training data and fails to generalize to test data, has become increasingly significant in the training of large neur…
Towards Efficient and Scalable Sharpness-Aware Minimization
Yong Liu, Siqi Mai, Xiangning Chen +2
Recently, Sharpness-Aware Minimization (SAM), which connects the geometry of the loss landscape and generalization, has demonstrated significant performance boosts on training larg…
RANK-NOSH: Efficient Predictor-Based Architecture Search via Non-Uniform Successive Halving
Ruochen Wang, Xiangning Chen, Minhao Cheng +2
Predictor-based algorithms have achieved remarkable performance in the Neural Architecture Search (NAS) tasks. However, these methods suffer from high computation costs, as trainin…
Rethinking Architecture Selection in Differentiable NAS
Ruochen Wang, Minhao Cheng, Xiangning Chen +2
Differentiable Neural Architecture Search is one of the most popular Neural Architecture Search (NAS) methods for its search efficiency and simplicity, accomplished by jointly opti…