CSS-LM: A Contrastive Framework for Semi-supervised Fine-tuning of Pre-trained Language Models
arXiv:2102.03752 · doi:10.1109/TASLP.2021.3105013
Abstract
Fine-tuning pre-trained language models (PLMs) has demonstrated its effectiveness on various downstream NLP tasks recently. However, in many low-resource scenarios, the conventional fine-tuning strategies cannot sufficiently capture the important semantic features for downstream tasks. To address this issue, we introduce a novel framework (named "CSS-LM") to improve the fine-tuning phase of PLMs via contrastive semi-supervised learning. Specifically, given a specific task, we retrieve positive and negative instances from large-scale unlabeled corpora according to their domain-level and class-level semantic relatedness to the task. We then perform contrastive semi-supervised learning on both the retrieved unlabeled and original labeled instances to help PLMs capture crucial task-related semantic features. The experimental results show that CSS-LM achieves better results than the conventional fine-tuning strategy on a series of downstream tasks with few-shot settings, and outperforms the latest supervised contrastive fine-tuning strategies. Our datasets and source code will be available to provide more details.
References in corpus (13)
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- ERNIE: Enhanced Representation through Knowledge Integration
- ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
- Contrastive Multiview Coding
- Towards a Human-like Open-Domain Chatbot
- CLEAR: Contrastive Learning for Sentence Representation
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- Revisiting Few-sample BERT Fine-tuning
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization
- BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
- A Mutual Information Maximization Perspective of Language Representation Learning
- Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning
- Self-training Improves Pre-training for Natural Language Understanding