activity
20212025
most citedEvaluating Generative Ad Hoc Information Retrieval

44 citations · 166 across the 21 of their papers we have counts for

collaborators
Showing cs.IRShow all

18 papers · 1 filter

cs.IR2025

AiReview: An Open Platform for Accelerating Systematic Reviews with LLMs

Xinyu Mao, Teerapong Leelanupab, Martin Potthast +2

Systematic reviews are fundamental to evidence-based medicine. Creating one is time-consuming and labour-intensive, mainly due to the need to screen, or assess, many studies for in…

cs.IR2025★ 2 cited

DenseReviewer: A Screening Prioritisation Tool for Systematic Review based on Dense Retrieval

Xinyu Mao, Teerapong Leelanupab, Harrisen Scells +1

Screening is a time-consuming and labour-intensive yet required task for medical systematic reviews, as tens of thousands of studies often need to be screened. Prioritising relevan…

cs.IR2025★ 3 cited

Variations in Relevance Judgments and the Shelf Life of Test Collections

Andrew Parry, Maik Fröbe, Harrisen Scells +5

The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditiona…

cs.IR2024

Ranking Generated Answers: On the Agreement of Retrieval Models with Humans on Consumer Health Questions

Sebastian Heineking, Jonas Probst, Daniel Steinbach +2

Evaluating the output of generative large language models (LLMs) is challenging and difficult to scale. Many evaluations of LLMs focus on tasks such as single-choice question-answe…

cs.IR2024★ 1 cited

Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins

Lukas Gienapp, Niklas Deckers, Martin Potthast +1

Representation-based retrieval models, so-called bi-encoders, estimate the relevance of a document to a query by calculating the similarity of their respective embeddings. Current…

cs.IR2024★ 8 cited

Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking

Ferdinand Schlatt, Maik Fröbe, Harrisen Scells +6

Cross-encoders distilled from large language models (LLMs) are often more effective re-rankers than cross-encoders fine-tuned on manually labeled data. However, distilled models do…