14 citations · 14 across the 5 of their papers we have counts for
5 papers
Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE
Qihuang Zhong, Liang Ding, Yibing Zhan +11
This technical report briefly describes our JDExplore d-team's Vega v2 submission on the SuperGLUE leaderboard. SuperGLUE is more challenging than the widely used general language…
Improving Sharpness-Aware Minimization with Fisher Mask for Better Generalization on Language Models
Qihuang Zhong, Liang Ding, Li Shen +4
Fine-tuning large pretrained language models on a limited training corpus usually suffers from poor generalization. Prior works show that the recently-proposed sharpness-aware mini…
Not All Instances Contribute Equally: Instance-adaptive Class Representation Learning for Few-Shot Visual Recognition
Mengya Han, Yibing Zhan, Yong Luo +4
Few-shot visual recognition refers to recognize novel visual concepts from a few labeled instances. Many few-shot visual recognition methods adopt the metric-based meta-learning pa…
Robust Weight Perturbation for Adversarial Training
Chaojian Yu, Bo Han, Mingming Gong +4
Overfitting widely exists in adversarial robust training of deep networks. An effective remedy is adversarial weight perturbation, which injects the worst-case weight perturbation…
Weakly-Supervised Semantic Segmentation with Visual Words Learning and Hybrid Pooling
Lixiang Ru, Bo Du, Yibing Zhan +1
Weakly-Supervised Semantic Segmentation (WSSS) methods with image-level labels generally train a classification network to generate the Class Activation Maps (CAMs) as the initial…