Artifact Sharing for Information Retrieval Research
arXiv:2505.05434 · doi:10.1145/3726302.3730147
Abstract
Sharing artifacts -- such as trained models, pre-built indexes, and the code to use them -- aids in reproducibility efforts by allowing researchers to validate intermediate steps and improves the sustainability of research by allowing multiple groups to build off one another's prior computational work. Although there are de facto consensuses on how to share research code (through a git repository linked to from publications) and trained models (via HuggingFace Hub), there is no consensus for other types of artifacts, such as built indexes. Given the practical utility of using shared indexes, researchers have resorted to self-hosting these resources or performing ad hoc file transfers upon request, ultimately limiting the artifacts' discoverability and reuse. This demonstration introduces a flexible and interoperable way to share artifacts for Information Retrieval research, improving both their accessibility and usability.
SIGIR 2025 (demo)
References in corpus (12)
- The Faiss library
- The Information Retrieval Experiment Platform
- Adaptive Re-Ranking with a Corpus Graph
- Lexically-Accelerated Dense Retrieval
- LexBoost: Improving Lexical Document Retrieval with Nearest Neighbors
- Neural Passage Quality Estimation for Static Pruning
- BM25S: Orders of magnitude faster lexical search via eager sparse scoring
- Quam: Adaptive Retrieval through Query Affinity Modelling
- Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
- Contextual Document Embeddings
- PyTerrier-GenRank: The PyTerrier Plugin for Reranking with Large Language Models
- Down with the Hierarchy: The 'H' in HNSW Stands for "Hubs"