5 citations · 7 across the 6 of their papers we have counts for
8 papers
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Lukas Gienapp, Christopher Schröder, Stefan Schweter +5
Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This prob…
Overview of the Plagiarism Detection Task at PAN 2025
André Greiner-Petter, Maik Fröbe, Jan Philip Wahle +4
The generative plagiarism detection task at PAN 2025 aims at identifying automatically generated textual plagiarism in scientific articles and aligning them with their respective s…
Investigating Counterclaims in Causality Extraction from Text
Tim Hagen, Niklas Deckers, Felix Wolter +2
Many causal claims, such as "sugar causes hyperactivity," are disputed or outdated. Yet research on causality extraction from text has almost entirely neglected counterclaims of ca…
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
Lukas Gienapp, Martin Potthast, Andrew Yates +2
The unjudged document problem, where systems that did not contribute to the original judgement pool may retrieve documents without a relevance judgement, is a key obstacle to the r…
Simplified Longitudinal Retrieval Experiments: A Case Study on Query Expansion and Document Boosting
Jüri Keller, Maik Fröbe, Gijs Hendriksen +3
The longitudinal evaluation of retrieval systems aims to capture how information needs and documents evolve over time. However, classical Cranfield-style retrieval evaluations only…
The Viability of Crowdsourcing for RAG Evaluation
Lukas Gienapp, Tim Hagen, Maik Fröbe +4
How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RA…