22 citations · 27 across the 7 of their papers we have counts for
6 papers · 1 filter
User Preference Induction with LLMs for Offline Top-N Recommendation Evaluation
David Otero, Javier Parapar
Offline evaluation is the standard methodology for comparing top-N recommender systems, yet it relies on incomplete relevance information. In most benchmark datasets, only a small…
Hybrid Pooling with LLMs via Relevance Context Learning
David Otero, Javier Parapar
High-quality relevance judgements over large query sets are essential for evaluating Information Retrieval (IR) systems, yet manual annotation remains costly and time-consuming. La…
LLM-Assisted Pseudo-Relevance Feedback
David Otero, Javier Parapar
Query expansion is a long-standing technique to mitigate vocabulary mismatch in ad hoc Information Retrieval. Pseudo-relevance feedback methods, such as RM3, estimate an expanded q…
Towards Reliable Testing for Multiple Information Retrieval System Comparisons
David Otero, Javier Parapar, Álvaro Barreiro
Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests…
Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation
David Otero, Javier Parapar, Álvaro Barreiro
Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating wh…
How Discriminative Are Your Qrels? How To Study the Statistical Significance of Document Adjudication Methods
David Otero, Javier Parapar, Nicola Ferro
Creating test collections for offline retrieval evaluation requires human effort to judge documents' relevance. This expensive activity motivated much work in developing methods fo…