Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
arXiv:2506.10802 · doi:10.1145/3726302.3730150
Abstract
Search engines often follow a pipeline architecture, where complex but effective reranking components are used to refine the results of an initial retrieval. Retrieval augmented generation (RAG) is an exciting application of the pipeline architecture, where the final component generates a coherent answer for the users from the retrieved documents. In this demo paper, we describe how such RAG pipelines can be formulated in the declarative PyTerrier architecture, and the advantages of doing so. Our PyTerrier-RAG extension for PyTerrier provides easy access to standard RAG datasets and evaluation measures, state-of-the-art LLM readers, and using PyTerrier's unique operator notation, easy-to-build pipelines. We demonstrate the succinctness of indexing and RAG pipelines on standard datasets (including Natural Questions) and how to build on the larger PyTerrier ecosystem with state-of-the-art sparse, learned-sparse, and dense retrievers, and other neural rankers.
4 pages, 3 tables, Accepted to SIGIR 2025
References in corpus (10)
- Retrieval-Augmented Generation for Large Language Models: A Survey
- The Power of Noise: Redefining Retrieval for RAG Systems
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Declarative Experimentation in Information Retrieval using PyTerrier
- Pseudo-Relevance Feedback for Multiple Representation Dense Retrieval
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
- Artifact Sharing for Information Retrieval Research
- PyTerrier-GenRank: The PyTerrier Plugin for Reranking with Large Language Models
- On Precomputation and Caching in Information Retrieval Experiments with Pipeline Architectures