163 citations · 197 across the 18 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022★ 13 cited
Surrogate Gap Minimization Improves Sharpness-Aware Training
Juntang Zhuang, Boqing Gong, Liangzhe Yuan +6
The recently proposed Sharpness-Aware Minimization (SAM) improves generalization by minimizing a \textit{perturbed loss} defined as the maximum loss within a neighborhood in the pa…
cs.CL2022
PERT: Pre-training BERT with Permuted Language Model
Yiming Cui, Ziqing Yang, Ting Liu
Pre-trained Language Models (PLMs) have been widely used in various natural language processing (NLP) tasks, owing to their powerful text representations trained on large-scale cor…