activity
20212024
most citedCorpusBrain: Pre-train a Generative Retrieval Model for Knowledge-Intensive Language Tasks

55 citations · 112 across the 14 of their papers we have counts for

collaborators

14 papers

cs.IR20242 cited

Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective

Yu-An Liu, Ruqing Zhang, Jiafeng Guo +3

Recent advances in neural information retrieval (IR) models have significantly enhanced their effectiveness over various IR tasks. The robustness of these models, essential for ens…

cs.CL2024

QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression

Wenshan Wang, Yihang Wang, Yixing Fan +2

In-context learning (ICL) capabilities are foundational to the success of large language models (LLMs). Recently, context compression has attracted growing interest since it can la…

cs.IR2024

Bootstrapped Pre-training with Dynamic Identifier Prediction for Generative Retrieval

Yubao Tang, Ruqing Zhang, Jiafeng Guo +3

Generative retrieval uses differentiable search indexes to directly generate relevant document identifiers in response to a query. Recent studies have highlighted the potential of…

cs.IR20241 cited

Multi-granular Adversarial Attacks against Black-box Neural Ranking Models

Yu-An Liu, Ruqing Zhang, Jiafeng Guo +3

Adversarial ranking attacks have gained increasing attention due to their success in probing vulnerabilities, and, hence, enhancing the robustness, of neural ranking models. Conven…

cs.IR20241 cited

CorpusBrain++: A Continual Generative Pre-Training Framework for Knowledge-Intensive Language Tasks

Jiafeng Guo, Changjiang Zhou, Ruqing Zhang +4

Knowledge-intensive language tasks (KILTs) typically require retrieving relevant documents from trustworthy corpora, e.g., Wikipedia, to produce specific answers. Very recently, a…

cs.IR2023

CAME: Competitively Learning a Mixture-of-Experts Model for First-stage Retrieval

Yinqiong Cai, Yixing Fan, Keping Bi +4

The first-stage retrieval aims to retrieve a subset of candidate documents from a huge collection both effectively and efficiently. Since various matching patterns can exist betwee…