Text Embeddings by Weakly-Supervised Contrastive Pre-training
arXiv:2212.03533
Abstract
This paper presents E5, a family of state-of-the-art text embeddings that transfer well to a wide range of tasks. The model is trained in a contrastive manner with weak supervision signals from our curated large-scale text pair dataset (called CCPairs). E5 can be readily used as a general-purpose embedding model for any tasks requiring a single-vector representation of texts such as retrieval, clustering, and classification, achieving strong performance in both zero-shot and fine-tuned settings. We conduct extensive evaluations on 56 datasets from the BEIR and MTEB benchmarks. For zero-shot settings, E5 is the first model that outperforms the strong BM25 baseline on the BEIR retrieval benchmark without using any labeled data. When fine-tuned, E5 obtains the best results on the MTEB benchmark, beating existing embedding models with 40x more parameters.
17 pages, v2 fixes the SummEval numbers
Cited by in corpus (12)
- Zero-shot Bilingual App Reviews Mining with Large Language Models
- Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
- Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
- Causal Question Answering with Reinforcement Learning
- EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning
- Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense Retrieval
- Word Sense Linking: Disambiguating Outside the Sandbox
- PEFA: Parameter-Free Adapters for Large-scale Embedding-based Retrieval Models
- Text Embedding Inversion Security for Multilingual Language Models
- SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
- Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval
- Vectorizing string entries for data processing on tables: when are larger language models better?