5 papers · 1 filter
Boolean queries are all you need?
Charles L. A. Clarke, Mark D. Smucker
We equipped an LLM-based search agent with access to a Boolean retrieval engine to search the MS MARCO V2.1 deduped segment collection used by the TREC 2024 RAG track. Over a stand…
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
Dake Zhang, Mark D. Smucker, Charles L. A. Clarke
Many readers today struggle to assess the trustworthiness of online news because reliable reporting coexists with misinformation. The TREC 2025 DRAGUN (Detection, Retrieval, and Au…
Extending MovieLens-32M to Provide New Evaluation Objectives
Mark D. Smucker, Houmaan Chamani
Offline evaluation of recommender systems has traditionally treated the problem as a machine learning problem. In the classic case of recommending movies, where the user has provid…
Assessing top- preferences
Charles L. A. Clarke, Alexandra Vtyurina, Mark D. Smucker
Assessors make preference judgments faster and more consistently than graded judgments. Preference judgments can also recognize distinctions between items that appear equivalent un…
Evaluating Sentence-Level Relevance Feedback for High-Recall Information Retrieval
Haotian Zhang, Gordon V. Cormack, Maura R. Grossman +1
This study uses a novel simulation framework to evaluate whether the time and effort necessary to achieve high recall using active learning is reduced by presenting the reviewer wi…