3 papers
cs.CL2025
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Lukas Gienapp, Christopher Schröder, Stefan Schweter +5
Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This prob…
cs.IR2025
Variations in Relevance Judgments and the Shelf Life of Test Collections
Andrew Parry, Maik Fröbe, Harrisen Scells +5
The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditiona…
cs.IR2024
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
Ferdinand Schlatt, Maik Fröbe, Matthias Hagen
A wide range of transformer-based language models have been proposed for information retrieval tasks. However, including transformer-based models in retrieval pipelines is often co…