How to Measure the Reproducibility of System-oriented IR Experiments
arXiv:2010.13447 · doi:10.1145/3397271.3401036
Abstract
Replicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towards more reproducible experimental practices and protocols, we also face a severe methodological issue: we do not have any means to assess when reproduced is reproduced. Moreover, we lack any reproducibility-oriented dataset, which would allow us to develop such methods. To address these issues, we compare several measures to objectively quantify to what extent we have replicated or reproduced a system-oriented IR experiment. These measures operate at different levels of granularity, from the fine-grained comparison of ranked lists, to the more general comparison of the obtained effects and significant differences. Moreover, we also develop a reproducibility-oriented dataset, which allows us to validate our measures and which can also be used to develop future measures.
SIGIR2020 Full Conference Paper
References in corpus (1)
Cited by in corpus (8)
- The Information Retrieval Experiment Platform
- How to Measure the Reproducibility of System-oriented IR Experiments
- A Reproducibility Study of PLAID
- repro_eval: A Python Interface to Reproducibility Measures of System-oriented IR Experiments
- Evaluation of Temporal Change in IR Test Collections
- LongEval at CLEF 2025: Longitudinal Evaluation of IR Systems on Web and Scientific Data
- Simplified Longitudinal Retrieval Experiments: A Case Study on Query Expansion and Document Boosting
- Evaluating Elements of Web-based Data Enrichment for Pseudo-Relevance Feedback Retrieval