Mask-guided BERT for Few Shot Text Classification
arXiv:2302.10447 · doi:10.1016/j.neucom.2024.128576
Abstract
Transformer-based language models have achieved significant success in various domains. However, the data-intensive nature of the transformer architecture requires much labeled data, which is challenging in low-resource scenarios (i.e., few-shot learning (FSL)). The main challenge of FSL is the difficulty of training robust models on small amounts of samples, which frequently leads to overfitting. Here we present Mask-BERT, a simple and modular framework to help BERT-based architectures tackle FSL. The proposed approach fundamentally differs from existing FSL strategies such as prompt tuning and meta-learning. The core idea is to selectively apply masks on text inputs and filter out irrelevant information, which guides the model to focus on discriminative tokens that influence prediction results. In addition, to make the text representations from different categories more separable and the text representations from the same category more compact, we introduce a contrastive learning loss function. Experimental results on public-domain benchmark datasets demonstrate the effectiveness of Mask-BERT.
References in corpus (16)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Learning from Few Examples: A Summary of Approaches to Few-Shot Learning
- The Falcon Series of Open Language Models
- Entailment as Few-Shot Learner
- AugGPT: Leveraging ChatGPT for Text Data Augmentation
- Meta-learning for Few-shot Natural Language Processing: A Survey
- Making LLMs Worth Every Penny: Resource-Limited Text Classification in Banking
- Few-Shot Text Classification with Triplet Networks, Data Augmentation, and Curriculum Learning
- Open, Closed, or Small Language Models for Text Classification?
- KNN-BERT: Fine-Tuning Pre-Trained Models with KNN Classifier
- Mask-guided Vision Transformer (MG-ViT) for Few-Shot Learning
- Few-shot learning for medical text: A systematic review
- CoT-Driven Framework for Short Text Classification: Enhancing and Transferring Capabilities from Large to Smaller Model
- On Evaluation Protocols for Data Augmentation in a Limited Data Scenario