6 papers
Beyond Relevance: Utility-Centric Retrieval in the LLM Era
Hengran Zhang, Minghao Tang, Keping Bi +1
Information retrieval systems have traditionally optimized for topical relevance-the degree to which retrieved documents match a query. However, relevance only approximates a deepe…
Annotation-Efficient Universal Honesty Alignment
Shiyu Ni, Keping Bi, Jiafeng Guo +4
Honesty alignment-the ability of large language models (LLMs) to recognize their knowledge boundaries and express calibrated confidence-is essential for trustworthy deployment. Exi…
Training Dense Retrievers with Multiple Positive Passages
Benben Wang, Minghao Tang, Hengran Zhang +2
Modern knowledge-intensive systems, such as retrieval-augmented generation (RAG), rely on effective retrievers to establish the performance ceiling for downstream modules. However,…
Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
Minghao Tang, Shiyu Ni, Jingtong Wu +2
Context-grounded generation underpins many LLM applications, including long-document question answering (QA), conversational personalization, and retrieval-augmented generation (RA…
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
Hengran Zhang, Minghao Tang, Keping Bi +5
This paper explores the use of large language models (LLMs) for annotating document utility in training retrieval and retrieval-augmented generation (RAG) systems, aiming to reduce…
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
Minghao Tang, Shiyu Ni, Jiafeng Guo +1
Retrieval-augmented generation (RAG) has been widely adopted to augment large language models (LLMs) with external knowledge for knowledge-intensive tasks. However, its effectivene…