5 citations · 5 across the 2 of their papers we have counts for
3 papers
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Lukas Gienapp, Christopher Schröder, Stefan Schweter +5
Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This prob…
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
Lukas Gienapp, Martin Potthast, Andrew Yates +2
The unjudged document problem, where systems that did not contribute to the original judgement pool may retrieve documents without a relevance judgement, is a key obstacle to the r…
The Viability of Crowdsourcing for RAG Evaluation
Lukas Gienapp, Tim Hagen, Maik Fröbe +4
How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RA…