Improving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives
arXiv:2206.02457 · doi:10.1145/3477495.3531823
Abstract
Following SimCSE, contrastive learning based methods have achieved the state-of-the-art (SOTA) performance in learning sentence embeddings. However, the unsupervised contrastive learning methods still lag far behind the supervised counterparts. We attribute this to the quality of positive and negative samples, and aim to improve both. Specifically, for positive samples, we propose switch-case augmentation to flip the case of the first letter of randomly selected words in a sentence. This is to counteract the intrinsic bias of pre-trained token embeddings to frequency, word cases and subwords. For negative samples, we sample hard negatives from the whole dataset based on a pre-trained language model. Combining the above two methods with SimCSE, our proposed Contrastive learning with Augmented and Retrieved Data for Sentence embedding (CARDS) method significantly surpasses the current SOTA on STS benchmarks in the unsupervised setting.
7 pages, 3 figures, 6 tables. Accepted to SIGIR 22. Code at https://github.com/alibaba/SimCSE-with-CARDS
References in corpus (12)
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- R-Drop: Regularized Dropout for Neural Networks
- Improving language models by retrieving from trillions of tokens
- Neural Text Generation with Unlikelihood Training
- Whitening Sentence Representations for Better Semantics and Faster Retrieval
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
- ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding
- Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval
- Incremental False Negative Detection for Contrastive Learning
- Frequency-based Distortions in Contextualized Word Embeddings
- Virtual Augmentation Supported Contrastive Learning of Sentence Representations