2 papers
cs.IR2025
Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation
David Otero, Javier Parapar, Ãlvaro Barreiro
Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating wh…
cs.IR2025
Towards Reliable Testing for Multiple Information Retrieval System Comparisons
David Otero, Javier Parapar, Ãlvaro Barreiro
Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests…