activity
20192026
most citedTowards General Text Embeddings with Multi-stage Contrastive Learning

66 citations · 107 across the 30 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.CL2023

Text Representation Distillation via Information Bottleneck Principle

Yanzhao Zhang, Dingkun Long, Zehan Li +1

Pre-trained language models (PLMs) have recently shown great success in text representation field. However, the high computational cost and high-dimensional representation of PLMs…

cs.IR2023

A Two-Stage Adaptation of Large Language Models for Text Ranking

Longhui Zhang, Yanzhao Zhang, Dingkun Long +3

Text ranking is a critical task in information retrieval. Recent advances in pre-trained language models (PLMs), especially large language models (LLMs), present new opportunities…

cs.CL2023

Language Models are Universal Embedders

Xin Zhang, Zehan Li, Yanzhao Zhang +4

In the large language model (LLM) revolution, embedding is a key component of various systems, such as retrieving knowledge or memories for LLMs or building content moderation filt…

cs.IR20231 cited

Hybrid Retrieval and Multi-stage Text Ranking Solution at TREC 2022 Deep Learning Track

Guangwei Xu, Yangzhao Zhang, Longhui Zhang +3

Large-scale text retrieval technology has been widely used in various practical business scenarios. This paper presents our systems for the TREC 2022 Deep Learning Track. We explai…

cs.CL202366 cited

Towards General Text Embeddings with Multi-stage Contrastive Learning

Zehan Li, Xin Zhang, Yanzhao Zhang +3

We present GTE, a general-purpose text embedding model trained with multi-stage contrastive learning. In line with recent advancements in unifying various NLP tasks into a single f…

cs.IR20231 cited

Challenging Decoder helps in Masked Auto-Encoder Pre-training for Dense Passage Retrieval

Zehan Li, Yanzhao Zhang, Dingkun Long +1

Recently, various studies have been directed towards exploring dense passage retrieval techniques employing pre-trained language models, among which the masked auto-encoder (MAE) p…