3 citations · 4 across the 6 of their papers we have counts for
6 papers · 1 filter
Benchmarking and Enabling Efficient Chinese Medical Retrieval via Asymmetric Encoders
Angqing Jiang, Jianlyu Chen, Zhe Fang +4
Effective medical text retrieval requires both high accuracy and low latency. While LLM-based embedding models possess powerful retrieval capabilities, their prohibitive latency an…
Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
Junwei Lan, Jianlyu Chen, Zheng Liu +3
With the growing popularity of LLM agents and RAG, it has become increasingly important to retrieve documents that are essential for solving a task, even when their connection to t…
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
Jianlyu Chen, Junwei Lan, Chaofan Li +2
In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes three key technical contributions. Fir…
Towards A Generalist Code Embedding Model Based On Massive Data Synthesis
Chaofan Li, Jianlyu Chen, Yingxia Shao +2
Code embedding models attract increasing attention due to the widespread popularity of retrieval-augmented generation (RAG) in software development. These models are expected to ca…
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
Jianlyu Chen, Nan Wang, Chaofan Li +6
Evaluation plays a crucial role in the advancement of information retrieval (IR) models. However, current benchmarks, which are based on predefined domains and human-labeled data,…
Making Text Embedders Few-Shot Learners
Chaofan Li, MingHao Qin, Shitao Xiao +5
Large language models (LLMs) with decoder-only architectures demonstrate remarkable in-context learning (ICL) capabilities. This feature enables them to effectively handle both fam…