activity
20192026
most citedTowards General Text Embeddings with Multi-stage Contrastive Learning

66 citations · 101 across the 25 of their papers we have counts for

collaborators
Showing 2024 · cs.CLShow all

6 papers · 2 filters

cs.CL2024

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Xin Zhang, Yanzhao Zhang, Wen Xie +7

Universal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a comb…

cs.CL2024

When Text Embedding Meets Large Language Model: A Comprehensive Survey

Zhijie Nie, Zhangchi Feng, Mingxin Li +4

Text embedding has become a foundational technology in natural language processing (NLP) during the deep learning era, driving advancements across a wide array of downstream tasks.…

cs.CL2024

Improving General Text Embedding Model: Tackling Task Conflict and Data Imbalance through Model Merging

Mingxin Li, Zhijie Nie, Yanzhao Zhang +3

Text embeddings are vital for tasks such as text retrieval and semantic textual similarity (STS). Recently, the advent of pretrained language models, along with unified benchmarks…

cs.CL2024

An End-to-End Model for Photo-Sharing Multi-modal Dialogue Generation

Peiming Guo, Sinuo Liu, Yanzhao Zhang +4

Photo-Sharing Multi-modal dialogue generation requires a dialogue agent not only to generate text responses but also to share photos at the proper moment. Using image text caption…

cs.CL2024

mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval

Xin Zhang, Yanzhao Zhang, Dingkun Long +10

We present systematic efforts in building long-context multilingual text representation model (TRM) and reranker from scratch for text retrieval. We first introduce a text encoder…

cs.CL20241 cited

Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-training

Longhui Zhang, Dingkun Long, Meishan Zhang +3

Chinese sequence labeling tasks are heavily reliant on accurate word boundary demarcation. Although current pre-trained language models (PLMs) have achieved substantial gains on th…