most citedMake Sharpness-Aware Minimization Stronger: A Sparsified Perturbation Approach

17 citations · 56 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL202214 cited

Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Qihuang Zhong, Liang Ding, Yibing Zhan +11

This technical report briefly describes our JDExplore d-team's Vega v2 submission on the SuperGLUE leaderboard. SuperGLUE is more challenging than the widely used general language…

cs.LG202217 cited

Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation Approach

Peng Mi, Li Shen, Tianhe Ren +4

Deep neural networks often suffer from poor generalization caused by complex and non-convex loss landscapes. One of the popular solutions is Sharpness-Aware Minimization (SAM), whi…

cs.CL20228 cited

On the Complementarity between Pre-Training and Random-Initialization for Resource-Rich Machine Translation

Changtong Zan, Liang Ding, Li Shen +3

Pre-Training (PT) of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains (someti…

cs.CL2022

Improving Sharpness-Aware Minimization with Fisher Mask for Better Generalization on Language Models

Qihuang Zhong, Liang Ding, Li Shen +4

Fine-tuning large pretrained language models on a limited training corpus usually suffers from poor generalization. Prior works show that the recently-proposed sharpness-aware mini…

cs.LG20224 cited

Robust Unlearnable Examples: Protecting Data Against Adversarial Learning

Shaopeng Fu, Fengxiang He, Yang Liu +2

The tremendous amount of accessible data in cyberspace face the risk of being unauthorized used for training deep learning models. To address this concern, methods are proposed to…

cs.LG202213 cited

Achieving Personalized Federated Learning with Sparse Local Models

Tiansheng Huang, Shiwei Liu, Li Shen +3

Federated learning (FL) is vulnerable to heterogeneously distributed data, since a common global model in FL may not adapt to the heterogeneous data distribution of each user. To c…