3 papers
cs.IR2026
Revisiting Text Ranking in Deep Research
Chuan Meng, Litu Ou, Sean MacAvaney +1
Deep research has emerged as an important task that aims to address hard queries that need extensive open-web exploration. To tackle it, most prior work equips large language model…
cs.IR2026
Re-Rankers as Relevance Judges
Chuan Meng, Jiqun Liu, Mohammad Aliannejadi +3
Using large language models (LLMs) to predict relevance judgments has shown promising results. Most studies treat this task as a distinct research line, e.g., focusing on prompt de…
cs.IR2025
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
Laura Dietz, Oleg Zendel, Peter Bailey +6
Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empi…