2 citations · 2 across the 3 of their papers we have counts for
3 papers
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
Maximilian Idahl, Benedikt Droste, Björn Plüster +1
Since FineWeb-Edu, data curation for LLM pretraining has predominantly relied on single scalar quality scores produced by small classifiers. A single score conflates multiple quali…
sui-1: Grounded and Verifiable Long-Form Summarization
Benedikt Droste, Jan Philipp Harries, Maximilian Idahl +1
Large language models frequently generate plausible but unfaithful summaries that users cannot verify against source text, a critical limitation in compliance-sensitive domains suc…
Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations
Björn Plüster, Jakob Ambsdorf, Lukas Braach +2
Natural language explanations promise to offer intuitively understandable explanations of a neural network's decision process in complex vision-language tasks, as pursued in recent…