collaborators

6 papers

cs.IR2026

Beyond Relevance: Utility-Centric Retrieval in the LLM Era

Hengran Zhang, Minghao Tang, Keping Bi +1

Information retrieval systems have traditionally optimized for topical relevance-the degree to which retrieved documents match a query. However, relevance only approximates a deepe…

cs.CL2026

Annotation-Efficient Universal Honesty Alignment

Shiyu Ni, Keping Bi, Jiafeng Guo +4

Honesty alignment-the ability of large language models (LLMs) to recognize their knowledge boundaries and express calibrated confidence-is essential for trustworthy deployment. Exi…

cs.IR2026

Training Dense Retrievers with Multiple Positive Passages

Benben Wang, Minghao Tang, Hengran Zhang +2

Modern knowledge-intensive systems, such as retrieval-augmented generation (RAG), rely on effective retrievers to establish the performance ceiling for downstream modules. However,…

cs.IR2026

Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation

Minghao Tang, Shiyu Ni, Jingtong Wu +2

Context-grounded generation underpins many LLM applications, including long-document question answering (QA), conversational personalization, and retrieval-augmented generation (RA…

cs.IR2025

Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation

Hengran Zhang, Minghao Tang, Keping Bi +5

This paper explores the use of large language models (LLMs) for annotating document utility in training retrieval and retrieval-augmented generation (RAG) systems, aiming to reduce…

cs.IR2025

Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation

Minghao Tang, Shiyu Ni, Jiafeng Guo +1

Retrieval-augmented generation (RAG) has been widely adopted to augment large language models (LLMs) with external knowledge for knowledge-intensive tasks. However, its effectivene…