5 citations · 24 across the 91 of their papers we have counts for
38 papers · 1 filter
Can LLM Rerankers Predict Their Own Ranking Performance?
Shiyu Ni, Keping Bi, Jiafeng Guo +3
Retrieval effectiveness varies substantially across queries, making it important to estimate ranking quality before relevance judgments are available. Query performance prediction…
AdversarialCoT: Single-Document Retrieval Poisoning for LLM Reasoning
Hongru Song, Yu-An Liu, Ruqing Zhang +4
Retrieval-augmented generation (RAG) enhances large language model (LLM) reasoning by retrieving external documents, but also opens up new attack surfaces. We study knowledge-base…
Data, Not Model: Explaining Bias toward LLM Texts in Neural Retrievers
Wei Huang, Keping Bi, Yinqiong Cai +3
Recent studies show that neural retrievers often display source bias, favoring passages generated by LLMs over human-written ones, even when both are semantically similar. This bia…
Bagging-Based Model Merging for Robust General Text Embeddings
Hengran Zhang, Keping Bi, Jiafeng Guo +4
General-purpose text embedding models underpin a wide range of NLP and information retrieval applications, and are typically trained on large-scale multi-task corpora to encourage…
Can LLM Annotations Replace User Clicks for Learning to Rank?
Lulu Yu, Keping Bi, Jiafeng Guo +4
Large-scale supervised data is essential for training modern ranking models, but obtaining high-quality human annotations is costly. Click data has been widely used as a low-cost a…
C2T-ID: Converting Semantic Codebooks to Textual Document Identifiers for Generative Search
Yingchen Zhang, Ruqing Zhang, Jiafeng Guo +4
Designing document identifiers (docids) that carry rich semantic information while maintaining tractable search spaces is a important challenge in generative retrieval (GR). Popula…