Identical and Fraternal Twins: Fine-Grained Semantic Contrastive Learning of Sentence Representations
arXiv:2307.10932 · doi:10.3233/FAIA230584
Abstract
The enhancement of unsupervised learning of sentence representations has been significantly achieved by the utility of contrastive learning. This approach clusters the augmented positive instance with the anchor instance to create a desired embedding space. However, relying solely on the contrastive objective can result in sub-optimal outcomes due to its inability to differentiate subtle semantic variations between positive pairs. Specifically, common data augmentation techniques frequently introduce semantic distortion, leading to a semantic margin between the positive pair. While the InfoNCE loss function overlooks the semantic margin and prioritizes similarity maximization between positive pairs during training, leading to the insensitive semantic comprehension ability of the trained model. In this paper, we introduce a novel Identical and Fraternal Twins of Contrastive Learning (named IFTCL) framework, capable of simultaneously adapting to various positive pairs generated by different augmentation techniques. We propose a \textit{Twins Loss} to preserve the innate margin during training and promote the potential of data enhancement in order to overcome the sub-optimal issue. We also present proof-of-concept experiments combined with the contrastive objective to prove the validity of the proposed Twins Loss. Furthermore, we propose a hippocampus queue mechanism to restore and reuse the negative instances without additional calculation, which further enhances the efficiency and performance of the IFCL. We verify the IFCL framework on nine semantic textual similarity tasks with both English and Chinese datasets, and the experimental results show that IFCL outperforms state-of-the-art methods.
This article has been accepted for publication in European Conference on Artificial Intelligence (ECAI2023). 9 pages, 4 figures
References in corpus (13)
- A Simple Framework for Contrastive Learning of Visual Representations
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- What Makes for Good Views for Contrastive Learning?
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- Whitening Sentence Representations for Better Semantics and Faster Retrieval
- How Language-Neutral is Multilingual BERT?
- CLUE: A Chinese Language Understanding Evaluation Benchmark
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer
- On the Sentence Embeddings from Pre-trained Language Models
- Virtual Augmentation Supported Contrastive Learning of Sentence Representations
- A Sentence is Worth 128 Pseudo Tokens: A Semantic-Aware Contrastive Learning Framework for Sentence Embeddings
- SNCSE: Contrastive Learning for Unsupervised Sentence Embedding with Soft Negative Samples
- Exploring the Impact of Negative Samples of Contrastive Learning: A Case Study of Sentence Embedding