most citedBrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent

2 citations · 2 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CL20252 cited

BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent

Zijian Chen, Xueguang Ma, Shengyao Zhuang +17

Deep-Research agents, which integrate large language models (LLMs) with search tools, have shown success in improving the effectiveness of handling complex queries that require ite…

cs.IR2025

MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed

Jiaqi Samantha Zhan, Crystina Zhang, Shengyao Zhuang +2

Effective video retrieval remains challenging due to the complexity of integrating visual, auditory, and textual modalities. In this paper, we explore unified retrieval methods usi…

cs.IR2025

Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Xueguang Ma, Luyu Gao, Shengyao Zhuang +3

Recent advancements in large language models (LLMs) have driven interest in billion-scale retrieval models with strong generalization across retrieval tasks and languages. Addition…

cs.IR2025

Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs

Nandan Thakur, Crystina Zhang, Xueguang Ma +1

Training robust retrieval and reranker models typically relies on large-scale retrieval datasets; for example, the BGE collection contains 1.6 million query-passage pairs sourced f…

cs.IR2025

Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning

Shengyao Zhuang, Xueguang Ma, Bevan Koopman +2

In this paper, we introduce Rank-R1, a novel LLM-based reranker that performs reasoning over both the user query and candidate documents before performing the ranking task. Existin…

cs.CL2025

DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers

Xueguang Ma, Xi Victoria Lin, Barlas Oguz +3

Large language models (LLMs) have demonstrated strong effectiveness and robustness while fine-tuned as dense retrievers. However, their large parameter size brings significant infe…