35 citations · 36 across the 6 of their papers we have counts for
5 papers · 1 filter
A Sovereign, Open-Source Foundation Model for German and English
Soofi-Team, :, Benedikt Droste +30
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B…
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
Maximilian Idahl, Benedikt Droste, Björn Plüster +1
Since FineWeb-Edu, data curation for LLM pretraining has predominantly relied on single scalar quality scores produced by small classifiers. A single score conflates multiple quali…
sui-1: Grounded and Verifiable Long-Form Summarization
Benedikt Droste, Jan Philipp Harries, Maximilian Idahl +1
Large language models frequently generate plausible but unfaithful summaries that users cannot verify against source text, a critical limitation in compliance-sensitive domains suc…
Multimodal Analytics for Real-world News using Measures of Cross-modal Entity Consistency
Eric Müller-Budack, Jonas Theiner, Sebastian Diering +2
The World Wide Web has become a popular source for gathering information and news. Multimodal information, e.g., enriching text with photos, is typically used to convey the news mo…