DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations
arXiv:2006.03659
Abstract
Sentence embeddings are an important component of many natural language processing (NLP) systems. Like word embeddings, sentence embeddings are typically learned on large text corpora and then transferred to various downstream tasks, such as clustering and retrieval. Unlike word embeddings, the highest performing solutions for learning sentence embeddings require labelled data, limiting their usefulness to languages and domains where labelled data is abundant. In this paper, we present DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations. Inspired by recent advances in deep metric learning (DML), we carefully design a self-supervised objective for learning universal sentence embeddings that does not require labelled training data. When used to extend the pretraining of transformer-based language models, our approach closes the performance gap between unsupervised and supervised pretraining for universal sentence encoders. Importantly, our experiments suggest that the quality of the learned embeddings scale with both the number of trainable parameters and the amount of unlabelled training data. Our code and pretrained models are publicly available and can be easily adapted to new domains or used to embed unseen text.
ACL2021 Camera Ready V2
References in corpus (7)
- Distributed Representations of Sentences and Documents
- Cross-lingual Language Model Pretraining
- Skip-Thought Vectors
- A Theoretical Analysis of Contrastive Unsupervised Representation Learning
- A Mutual Information Maximization Perspective of Language Representation Learning
- PyTorch Metric Learning
- Cross-Batch Memory for Embedding Learning
Cited by in corpus (32)
- CLEAR: Contrastive Learning for Sentence Representation
- Contrastive Code Representation Learning
- Contrastive Self-supervised Sequential Recommendation with Robust Augmentation
- SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer
- TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning
- Towards the Generalization of Contrastive Self-Supervised Learning
- Contrastive encoder pre-training-based clustered federated learning for heterogeneous data
- Investigating the Role of Negatives in Contrastive Representation Learning
- Self-Guided Contrastive Learning for BERT Sentence Representations
- Supporting Clustering with Contrastive Learning
- Adversarial Training with Contrastive Learning in NLP
- Contrastive Learning for Robust Android Malware Familial Classification
- Interest-oriented Universal User Representation via Contrastive Learning
- Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup
- Cross-Domain Sentiment Classification with In-Domain Contrastive Learning
- Unsupervised Document Embedding via Contrastive Augmentation
- Contrastive Attraction and Contrastive Repulsion for Representation Learning
- Disentangled Contrastive Learning for Learning Robust Textual Representations
- Bi-Granularity Contrastive Learning for Post-Training in Few-Shot Scene
- Cross-Domain Sentiment Classification with Contrastive Learning and Mutual Information Maximization
- DialogueCSE: Dialogue-based Contrastive Learning of Sentence Embeddings
- PAUSE: Positive and Annealed Unlabeled Sentence Embedding
- Data-Efficient Language-Supervised Zero-Shot Learning with Self-Distillation
- Pairwise Supervised Contrastive Learning of Sentence Representations
- Contrastive String Representation Learning using Synthetic Data
- Lightweight Cross-Lingual Sentence Representation Learning
- Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast
- Contrastive Representation Learning for Exemplar-Guided Paraphrase Generation
- Exploiting Twitter as Source of Large Corpora of Weakly Similar Pairs for Semantic Sentence Embeddings
- Contrastive Document Representation Learning with Graph Attention Networks
- Are Classes Clusters?