activity
20232025
most citedSystematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR

9 citations · 18 across the 5 of their papers we have counts for

collaborators

9 papers

cs.IR20255 cited

The Viability of Crowdsourcing for RAG Evaluation

Lukas Gienapp, Tim Hagen, Maik Fröbe +4

How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RA…

cs.IR2025

Counterfactual Query Rewriting to Use Historical Relevance Feedback

Jüri Keller, Maik Fröbe, Gijs Hendriksen +4

When a retrieval system receives a query it has encountered before, previous relevance feedback, such as clicks or explicit judgments can help to improve retrieval results. However…

cs.IR2024

Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval

Ferdinand Schlatt, Maik Fröbe, Matthias Hagen

A wide range of transformer-based language models have been proposed for information retrieval tasks. However, including transformer-based models in retrieval pipelines is often co…

cs.IR20249 cited

Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR

Nandan Thakur, Luiz Bonifacio, Maik Fröbe +5

The zero-shot effectiveness of neural retrieval models is often evaluated on the BEIR benchmark -- a combination of different IR evaluation datasets. Interestingly, previous studie…

cs.CL20243 cited

Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance

Kaitlyn Zhou, Jena D. Hwang, Xiang Ren +3

The ability to communicate uncertainty, risk, and limitation is crucial for the safety of large language models. However, current evaluations of these abilities rely on simple cali…

cs.IR2024

Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking

Ferdinand Schlatt, Maik Fröbe, Harrisen Scells +6

Cross-encoders distilled from large language models (LLMs) are often more effective re-rankers than cross-encoders fine-tuned on manually labeled data. However, distilled models do…