7 papers
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
Jüri Keller, Maik Fröbe, Björn Engelmann +4
Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In e…
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation
Xuhong He, To Eun Kim, Maik Fröbe +3
Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingu…
Overview of PAN 2026: Voight-Kampff Generative AI Detection, Text Watermarking, Multi-Author Writing Style Analysis, Generative Plagiarism Detection, and Reasoning Trajectory Detection
Janek Bevendorff, Maik Fröbe, André Greiner-Petter +9
The goal of the PAN workshop is to advance computational stylometry and text forensics via objective and reproducible evaluation. In 2026, we run the following five tasks: (1) Voig…
Overview of the Plagiarism Detection Task at PAN 2025
André Greiner-Petter, Maik Fröbe, Jan Philip Wahle +4
The generative plagiarism detection task at PAN 2025 aims at identifying automatically generated textual plagiarism in scientific articles and aligning them with their respective s…
Variations in Relevance Judgments and the Shelf Life of Test Collections
Andrew Parry, Maik Fröbe, Harrisen Scells +5
The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditiona…
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking
Ferdinand Schlatt, Maik Fröbe, Harrisen Scells +6
Cross-encoders distilled from large language models (LLMs) are often more effective re-rankers than cross-encoders fine-tuned on manually labeled data. However, distilled models do…